nowfound

Alternatives

Products that do what Production grade end to end open source stack for Voice AI does

we have been building an open source orchestration which enables you to plug in your own TTS/ASR/LLM for end-to-end voice conversations at -> https://github.com/bolna-ai/bolna. Few days back, was having a discussion here in HN about the possibilities of having a complete open source stack for ASR+LLM+TTS. Today, we are releasing a complete open sourced Dockerized stack by merging Bolna with Whisper ASR, Llama3 and Melo TTS.

  1. 1IB

    I built a voice agent from scratch that averages ~400ms end-to-end latency (phone stop → first syllable). That’s with full STT → LLM → TTS in the loop, clean barge-ins, and no precomputed responses. What moved the needle: Voice is a turn-taking problem, not a transcription problem. VAD alone fails; you need semantic end-of-turn detection. The system reduces to one loop: speaking vs listening. The two transitions - cancel instantly on barge-in, respond instantly on end-of-turn - define the experience. STT → LLM → TTS must stream. Sequential pipelines are dead on arrival for natural…

    Mar 2026 · ntik.me

  2. 2TT
  3. 3

    Build Powerful Voice Agents

    2025

  4. 4AO

    I've been obsessed for the past ~year with the possibilities of talking to LLMs. I built a bunch of one-off prototypes, shared code on X, started a Meetup group in SF, and co-hosted a big hackathon. It turns out that there are a few low-level problems that everybody building conversational/real-time AI needs to solve on the way to building/shipping something that works well: low-latency media transport, echo cancellation, voice activity detection, phrase endpointing, pipelining data between models/services, handling voice interruptions, swapping out different…

    2024 · github.com

  5. 5IO

    Hi HN! Last year the project I launched here got a lot of good feedback on creating speech to speech AI on the ESP32. Recently I revamped the whole stack, iterated on that feedback and made our project fully open-source—all of the client, hardware, firmware code. This Github repo turns an ESP32-S3 into a realtime AI speech companion using the OpenAI Realtime API, Arduino WebSockets, Deno Edge Functions, and a full-stack web interface. You can talk to your own custom AI character, and it responds instantly. I couldn't find a resource that helped set up a reliable, secure websocket (WSS) AI…

    2025 · github.com

  6. 6IM

    A few years ago, right after high school, I decided to try to make a simultaneous translation app for Android as a side project, it took longer than expected (about 2 years) and I had to make a lot of compromises (I had to use Google's API and therefore make users use a developer key because at the time there were no free solutions for speech recognition and translation that had good quality). At the end of university, I decided to pick it up again and finally, using OpenAi's Whisper for speech recognition and Meta's NLLB for translation (with both running locally on the phone), I managed to…

    2024 · github.com

  7. 7

    Open-source TTS with emotion & voice cloning

    2025

  8. 8MO

    I wanted to share our new speech to text model, and the library to use them effectively. We're a small startup (six people, sub-$100k monthly GPU budget) so I'm proud of the work the team has done to create streaming STT models with lower word-error rates than OpenAI's largest Whisper model. Admittedly Large v3 is a couple of years old, but we're near the top the HF OpenASR leaderboard, even up against Nvidia's Parakeet family. Anyway, I'd love to get feedback on the models and software, and hear about what people might build with it.

    Feb 2026 · github.com

  9. 9
    OpenWispr190

    100% local open source AI speech-to-text model

    2025

  10. 10LV
  11. 11PS

    Welcome to Project S.A.T.U.R.D.A.Y. This is a project that allows anyone to easily build their own self-hosted J.A.R.V.I.S-like voice assistant. In my mind vocal computing is the future of human-computer interaction and by open sourcing this code I hope to expedite us on that path. I have had a blast working on this so far and I'm excited to continue to build with it. It uses whisper.cpp [1], Coqui TTS [2] and OpenAI [3] to do speech-to-text, text-to-text and text-to-speech inference all 100% locally (except for text-to-text). In the future I plan to swap out OpenAI for llama.cpp [4]. It is…

    2023 · github.com

  12. 12
    MARS5 TTS489

    Open-source, insanely prosodic text-to-speech model

    2024

  13. 13
    Bolna215

    Voice AI agents platform for high volume recruitment & ATS

    2025

  14. 14WL

    WhisperFusion builds upon the capabilities of open source tools WhisperLive and WhisperSpeech to provide a seamless conversations with an AI chatbot.

    2024 · github.com

  15. 15OS

    Our goal with this project is to build a completely open source, state of the art turn detection model that can be used in any voice AI application. I've been experimenting with LLM voice conversations since GPT-4 was first released. (There's a previous front page Show HN about Pipecat, the open source voice AI orchestration framework I work on. [1]) It's been almost two years, and for most of that time, I've been expecting that someone would "solve" turn detection. We all built initial, pretty good 80/20 versions of turn detection on top of VAD (voice activity detection) models. And…

    2025 · github.com

  16. 16OV

    Hi Hackernews, we're Maitreya, Prateek and Marmik. Over the past few months we've been working on building a platform to build, scale and monitor voice based LLM applications. Demo (https://www.youtube.com/watch?v=OSrOmyR7oQs) 1⃣ Open Source orchestration: We're open-sourcing our orchestration to quickly setup and create LLM based voice driven conversational applications https://github.com/bolna-ai/bolna/ 2⃣ Hosted API Platform: Exposing our managed solution via APIs to build voice driven applications…

    2024 · bolna.dev

  17. 17OS

    Hey HN, we've been working with OpenAI for the past few months on the new Realtime API. The goal is to give everyone access to the same stack that underpins Advanced Voice in the ChatGPT app. Under the hood it works like this: - A user's speech is captured by a LiveKit client SDK in the ChatGPT app - Their speech is streamed using WebRTC to OpenAI’s voice agent - The agent relays the speech prompt over websocket to GPT-4o - GPT-4o runs inference and streams speech packets (over websocket) back to the agent - The agent relays generated speech using WebRTC back to the user’s device The…

    2024 · github.com

  18. 18BB

    Hi Hacker News! This is Maitreya, Marmik and Prateek, co-founders of Bolna (https://github.com/bolna-ai/bolna). With Bolna, developers can create end-to-end conversational voice agents. They can connect to their own custom LLMs, their own Telephony, their own models etc. and create application features requiring voice AI. Here’s a small video: https://github.com/bolna-ai/bolna/assets/1313096/2237f64f-1c.... Our product originates from building an AI interviewer bot which can be used for practising coding interviews like Leetcode. By…

    2024 · github.com

  19. 19

    The voice for your real-time AI applications

    2025

  20. 20
    Qwen3-TTS155

    Voice design, cloning & 97ms streaming

    Jan 2026 · qwen.ai

  21. 21

    AI-powered speech to text, tts, & translation Mac app

    Oct 2025

  22. 22

    Open-Source Multilingual Text-to-Speech with Voice Cloning

    2024

  23. 23VA

    Voxos is an open-source desktop voice assistant that aims to put Clippy to shame while supporting new desktop workflows powered by LLMs. Tired of copy and pasting ChatGPT responses between your web browser and IDE? Does your copilot not quite do what you need it to do? I invite you to give Voxos a try and maybe even become a contributor!

    2024 · gitlab.com

  24. 24

    Fast, accurate STT and TTS APIs at the best price

    Apr 2026 · x.ai

Ranked by how close each launch is in meaning, then by votes. Refine with a description →