nowfound

AI · alternatives · 2026

24 alternatives to Silero VAD

One voice detector to rule them all

Below are 24 products that do a similar job, ranked by how close each is in meaning and then by launch-day votes. Silero VAD launched in 2021; newer entries below may have overtaken it.

  1. 1IB

    I built a voice agent from scratch that averages ~400ms end-to-end latency (phone stop → first syllable). That’s with full STT → LLM → TTS in the loop, clean barge-ins, and no precomputed responses. What moved the needle: Voice is a turn-taking problem, not a transcription problem. VAD alone fails; you need semantic end-of-turn detection. The system reduces to one loop: speaking vs listening. The two transitions - cancel instantly on barge-in, respond instantly on end-of-turn - define the experience. STT → LLM → TTS must stream. Sequential pipelines are dead on arrival for natural…

    Mar 2026 · ntik.me · its alternatives →

  2. 2KT

    Kitten TTS is an open-source series of tiny and expressive text-to-speech models for on-device applications. We are excited to launch a preview of our smallest model, which is less than 25 MB. This model has 15M parameters. This release supports English text-to-speech applications in eight voices: four male and four female. The model is quantized to int8 + fp16, and it uses onnx for runtime. The model is designed to run literally anywhere eg. raspberry pi, low-end smartphones, wearables, browsers etc. No GPU required! We're releasing this to give early users a sense of the latency and voices…

    2025 · github.com · its alternatives →

  3. 3TN

    Kitten TTS (https:&#x2F;&#x2F;github.com&#x2F;KittenML&#x2F;KittenTTS) is an open-source series of tiny and expressive text-to-speech models for on-device applications. We had a thread last year here: https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=44807868. Today we're releasing three new models with 80M, 40M and 14M parameters. The largest model (80M) has the highest quality. The 14M variant reaches new SOTA in expressivity among similar sized models, despite being <25MB in size. This release is a major upgrade from the previous one and supports English text-to-speech applications in…

    Mar 2026 · github.com · its alternatives →

  4. 4
    Vapi▲498

    Voice AI infrastructure for the internet

    2024 · its alternatives →

  5. 5
    Voxtral▲201

    Frontier open source speech understanding models

    2025 · mistral.ai · its alternatives →

  6. 6

    Build Powerful Voice Agents

    2025 · its alternatives →

  7. 7MO

    I wanted to share our new speech to text model, and the library to use them effectively. We're a small startup (six people, sub-$100k monthly GPU budget) so I'm proud of the work the team has done to create streaming STT models with lower word-error rates than OpenAI's largest Whisper model. Admittedly Large v3 is a couple of years old, but we're near the top the HF OpenASR leaderboard, even up against Nvidia's Parakeet family. Anyway, I'd love to get feedback on the models and software, and hear about what people might build with it.

    Feb 2026 · github.com · its alternatives →

  8. 8BY

    Voice Activity Detection (VAD) is a crucial component for Voice AI, enabling more natural and efficient interactions. TEN VAD is an open-source solution designed to supercharge your Voice AI Agents with lightning-fast, human-like conversations! TEN VAD offers some key advantages: ONNX Support: Deploy on virtually any platform or hardware architecture! This means greater flexibility and easier integration into your existing systems. Superior Detection Accuracy: Experience noticeable improvements in voice detection, leading to fewer errors and more reliable performance. Smaller & Faster: Enjoy…

    2025 · github.com · its alternatives →

  9. 9

    Highly expressive and natural speech generation model

    2025 · microsoft.ai · its alternatives →

  10. 10

    Production ASR for noisy multilingual audio

    Apr 2026 · microsoft.ai · its alternatives →

  11. 11

    Multilingual TTS model with realistic and expressive speech

    Mar 2026 · mistral.ai · its alternatives →

  12. 12OS

    Our goal with this project is to build a completely open source, state of the art turn detection model that can be used in any voice AI application. I've been experimenting with LLM voice conversations since GPT-4 was first released. (There's a previous front page Show HN about Pipecat, the open source voice AI orchestration framework I work on. [1]) It's been almost two years, and for most of that time, I've been expecting that someone would "solve" turn detection. We all built initial, pretty good 80&#x2F;20 versions of turn detection on top of VAD (voice activity detection) models. And…

    2025 · github.com · its alternatives →

  13. 13

    A neural net for speech recognition

    2022 · its alternatives →

  14. 14
    Vois 2.0▲105

    The ElevenLabs alternative with unlimited generation

    21d ago · vois.so · its alternatives →

  15. 15
    Voicr▲254

    Your voice in, polished text out — in seconds

    Mar 2026 · voicr.pro · its alternatives →

  16. 16

    Ultra-realistic text-to-speech

    2025 · vogent.ai · its alternatives →

  17. 17

    Next-Gen Speech-to-Text with Unmatched Performance

    2023 · its alternatives →

  18. 18
    Vapi CLI▲160

    The best DX for building voice AI

    2025 · vapi.ai · its alternatives →

  19. 19
    VoxCPM2▲110

    Open-source 48kHz TTS with voice design and cloning

    Apr 2026 · github.com · its alternatives →

  20. 20
    Voysis▲111

    The complete independent Voice AI platform

    2018 · its alternatives →

  21. 21
    Vocol.AI▲140

    All-in-one voice collaboration platform

    2023 · its alternatives →

  22. 22VA

    Voxos is an open-source desktop voice assistant that aims to put Clippy to shame while supporting new desktop workflows powered by LLMs. Tired of copy and pasting ChatGPT responses between your web browser and IDE? Does your copilot not quite do what you need it to do? I invite you to give Voxos a try and maybe even become a contributor!

    2024 · gitlab.com · its alternatives →

  23. 23AO
  24. 24
    HiDock H1▲118

    GPT powered audio dock with AI summary

    2023 · its alternatives →

Also compare

Ranked by how close each launch is in meaning, then by votes. Prices were read from each product’s own site when checked and can change. Refine with your own description →