Alternatives
Products that do what VoiceFlow does
Practice speaking. Kill filler words. In real time.
- 1

- 2

- 3IB
I built a voice agent from scratch that averages ~400ms end-to-end latency (phone stop → first syllable). That’s with full STT → LLM → TTS in the loop, clean barge-ins, and no precomputed responses. What moved the needle: Voice is a turn-taking problem, not a transcription problem. VAD alone fails; you need semantic end-of-turn detection. The system reduces to one loop: speaking vs listening. The two transitions - cancel instantly on barge-in, respond instantly on end-of-turn - define the experience. STT → LLM → TTS must stream. Sequential pipelines are dead on arrival for natural…
Mar 2026 · ntik.me
- 4

- 5

- 6IR
Hey HN, this is Lina, Andrew, and Sidney from Infinity AI (https://infinity.ai/). We've trained our own foundation video model focused on people. As far as we know, this is the first time someone has trained a video diffusion transformer that’s driven by audio input. This is cool because it allows for expressive, realistic-looking characters that actually speak. Here’s a blog with a bunch of examples: https://toinfinityai.github.io/v2-launch-page/ If you want to try it out, you can either (1) go to https://studio.infinity.ai/try-inf2, or (2)…
2024
- 7

- 8

- 9

- 10

- 11

- 12SC
2016 · itunes.apple.com
- 13

- 14EV
Our team is filled with technologists and creators, and when we record and edit videos, 80% of the time is spent chopping up the video, removing silences, and picking the right takes. So we decided to build a tool that did that for you — or at least get you there most of the way! Our initial implementation is somewhat naïve and uses a user configurable silence threshold that just reads in volume levels. In the future, we’d like to use a frequency-based approach that focuses on the human voice. We’re also open to ideas, so let us know if you have any!
2022 · kapwing.com
- 15

The most accurate streaming speech model for voice agents.
Mar 2026
- 16

- 17

- 18

- 19

- 20

- 21

- 22
- 23AP
TLDR: We created a personalised Andrej Karpathy tutor that can response to questions about his Youtube videos in sub 1 second responses (voice-to-voice). We do this using a voice enabled RAG agent. See later in the post for demo link, Github Repo and blog write up. A few weeks ago we released the worlds fastest voice bot, achieving 500ms voice-to-voice response times, including a 200ms delay waiting for a user to stop speaking. After reaching the front page of HN, we thought about how we could take this a step further based on feedback we were getting from the community. Many companies were…
2024 · educationbot.cerebrium.ai
- 24

Ranked by how close each launch is in meaning, then by votes. Refine with a description →