Alternatives
Products that do what Real-time voice chat with AI, no transcription does
Hi HN -- voice chat with AI is very popular these days, especially with YC startups (https://twitter.com/k7agar/status/1769078697661804795). The current approaches all do a cascaded approach, with audio -> transcription -> language model -> text synthesis. This approach is easy to get started with, but requires lots of complexity and has a few glaring limitations. Most notably, transcription is slow, is lossy and any error propagates to the rest of the system, cannot capture emotional affect, is often not robust to code-switching/accents, and more. Instead, what…
- 1RT
2025 · github.com
- 2

- 3

- 4

- 5

- 6IB
Hi HN, I built this out of frustration of the evergrowing list of AI models and features to try and to fit my workflow. The visual approach clicks for me so i went with it, it provides more freedom and control of the outcome, because predictable results and increased productivity is what I’m after when using conversational AI. The app is packed with features, my most used are prompt library, voice input and text search, narration is useful too. The app is local-first and works right in the browser, no sign up needed and it's absolutely free to try. BYOAK – bring your own API Keys. Let me…
2024 · grafychat.com
- 7

Powering the next-gen of smart, trusted voice agents
2025 · elevenlabs.io
- 8IR
Hey HN, this is Lina, Andrew, and Sidney from Infinity AI (https://infinity.ai/). We've trained our own foundation video model focused on people. As far as we know, this is the first time someone has trained a video diffusion transformer that’s driven by audio input. This is cool because it allows for expressive, realistic-looking characters that actually speak. Here’s a blog with a bunch of examples: https://toinfinityai.github.io/v2-launch-page/ If you want to try it out, you can either (1) go to https://studio.infinity.ai/try-inf2, or (2)…
2024
- 9

- 10WL
WhisperFusion builds upon the capabilities of open source tools WhisperLive and WhisperSpeech to provide a seamless conversations with an AI chatbot.
2024 · github.com
- 11

- 12

- 13

- 14

- 15
- 16IM
2022 · freesubtitles.ai
- 17VB
Last year when GPT-4 was released I started making lots of little voice + LLM experiments. Voice interfaces are fun; there are several interesting new problem spaces to explore. I'm convinced that voice is going to be a bigger and bigger part of how we all interact with generative AI. But one thing that's hard, today, is building voice bots that respond as quickly as humans do in conversation. A 500ms voice-to-voice response time is just barely possible with today's AI models. You can get down to 500ms if you: host transcription, LLM inference, and voice generation all together in one place;…
2024 · fastvoiceagent.cerebrium.ai
- 18

- 19

- 20

- 21

- 22

- 23

- 24IB
Hi, I built TalkBits because most language apps focus on vocabulary or exercises, but not actual conversation. The hard part of learning a language is speaking naturally under pressure. TalkBits lets you have real-time spoken conversations with an AI that acts like a native speaker. You can choose different scenarios (travel, daily life, work, etc.), speak naturally, and the AI responds with natural speech back. The goal is to make it feel like talking to a real person rather than doing lessons. Techwise, it uses realtime speech input, transcription, LLM responses, and tts streaming to keep…
Jan 2026 · apps.apple.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →