nowfound

Alternatives

Products that do what Voice gender classifier for European voice AI (1MB, ONNX, 4ms) does

Hi, I'm Kamil and I'm a founder of Applied AI agency in Warsaw, Poland. We've trained a small <1MB voice classifier model that runs on CPU in 4ms. Can be run next to silero VAD in voice AI deployments. What we noticed in production deployments of voice assistants in Contact Centers in EU is that human consultants pick up immediately how to inflect verbs and ajdectives after one utterance from the caller. But voice AI agents don't know it until 1-2 minutes into the call when they are either corrected or the caller uses explicitly words with male&#x2F;female form a couple of times. Our model…

  1. 1RT
  2. 2

    Multilingual speech AI model trained on 12.5M hours of data

    2024

  3. 3
    Cols.ai139

    AI phone calling platform

    2024

  4. 4VB

    Last year when GPT-4 was released I started making lots of little voice + LLM experiments. Voice interfaces are fun; there are several interesting new problem spaces to explore. I'm convinced that voice is going to be a bigger and bigger part of how we all interact with generative AI. But one thing that's hard, today, is building voice bots that respond as quickly as humans do in conversation. A 500ms voice-to-voice response time is just barely possible with today's AI models. You can get down to 500ms if you: host transcription, LLM inference, and voice generation all together in one place;…

    2024 · fastvoiceagent.cerebrium.ai

  5. 5

    Premium AI voice quality without the premium price tag.

    2025

  6. 6

    Highly expressive and natural speech generation model

    2025

  7. 7IB

    I built a voice agent from scratch that averages ~400ms end-to-end latency (phone stop → first syllable). That’s with full STT → LLM → TTS in the loop, clean barge-ins, and no precomputed responses. What moved the needle: Voice is a turn-taking problem, not a transcription problem. VAD alone fails; you need semantic end-of-turn detection. The system reduces to one loop: speaking vs listening. The two transitions - cancel instantly on barge-in, respond instantly on end-of-turn - define the experience. STT → LLM → TTS must stream. Sequential pipelines are dead on arrival for natural…

    Mar 2026 · ntik.me

  8. 8AP

    TLDR: We created a personalised Andrej Karpathy tutor that can response to questions about his Youtube videos in sub 1 second responses (voice-to-voice). We do this using a voice enabled RAG agent. See later in the post for demo link, Github Repo and blog write up. A few weeks ago we released the worlds fastest voice bot, achieving 500ms voice-to-voice response times, including a 200ms delay waiting for a user to stop speaking. After reaching the front page of HN, we thought about how we could take this a step further based on feedback we were getting from the community. Many companies were…

    2024 · educationbot.cerebrium.ai

  9. 9

    Generative AI voice bot for sales, support, and beyond

    2023

  10. 10

    First TTS model to support all 22 Indic languages + English

    2024

  11. 11

    Learn English, Spanish, German, French & more by talking

    2024

  12. 12

    Tighter instruction adherence in speech agents

    Feb 2026

  13. 13

    The voice-first AI assistant that takes action

    2025

  14. 14

    Voice AI that’s 5% of the cost. 100% of the quality.

    2025

  15. 15

    Human-like voices for every content

    2023

  16. 16
    Callin.io125

    The first AI phone assistant for small businesses

    2024

  17. 17

    AI Voice assistant in 3 minutes. Built for non-developers

    Nov 2025

  18. 18

    Analytics for your voice AI agent

    2024

  19. 19

    One API to build production-ready voice agents

    Apr 2026 · assemblyai.com

  20. 20

    Multilingual TTS model with realistic and expressive speech

    Mar 2026

  21. 21

    High-quality text-to-speech, designed for developers

    2025

  22. 22

    Voice AI that feels as good as it sounds

    May 2026 · inworld.ai

  23. 23OS

    Our goal with this project is to build a completely open source, state of the art turn detection model that can be used in any voice AI application. I've been experimenting with LLM voice conversations since GPT-4 was first released. (There's a previous front page Show HN about Pipecat, the open source voice AI orchestration framework I work on. [1]) It's been almost two years, and for most of that time, I've been expecting that someone would "solve" turn detection. We all built initial, pretty good 80&#x2F;20 versions of turn detection on top of VAD (voice activity detection) models. And…

    2025 · github.com

  24. 24
    Articula171

    Speak any language with your own voice

    2024

Ranked by how close each launch is in meaning, then by votes. Refine with a description →