Alternatives
Products that do what SpeakStruct – Turn voice into consistent structured data does
Hey folks, Built SpeakStruct to allow users to setup templates to turn voice input into consistent, structured output. Use cases from feedback I've had are customer support, coaching/check-ins, note taking, etc. Although there is a pricing section, signing up is free (no CC required). If you don't want to sign up, a demo is available here (sale-sy demo, but shows the product). https://app.arcade.software/share/nWm35szNPwD3PpH4eUSp Open to all feedback.
- 1
- 2WO
Hey HN! Braden here, creator of Whispering, an open-source speech-to-text app. I really like dictation. For years, I relied on transcription tools that were almost good, but they were all closed-source. Even a lot of them that claimed to be “local” or “on-device” were still black boxes that left me wondering where my audio really went. So I built Whispering. It’s open-source, local-first, and most importantly, transparent with your data. Your data is stored locally on your device, and your audio goes directly from your machine to a local provider (Whisper C++, Speaches, etc.) or your chosen…
2025 · github.com
- 3

- 4IB
I built a voice agent from scratch that averages ~400ms end-to-end latency (phone stop → first syllable). That’s with full STT → LLM → TTS in the loop, clean barge-ins, and no precomputed responses. What moved the needle: Voice is a turn-taking problem, not a transcription problem. VAD alone fails; you need semantic end-of-turn detection. The system reduces to one loop: speaking vs listening. The two transitions - cancel instantly on barge-in, respond instantly on end-of-turn - define the experience. STT → LLM → TTS must stream. Sequential pipelines are dead on arrival for natural…
Mar 2026 · ntik.me
- 5

- 6

- 7
- 8

- 9

Expressive cloned voices with text-to-speech in Unity
2020
- 10

- 11

- 12

- 13

- 14VP
Imagine creating a podcast where Mark Zuckerberg interviews Elon Musk – using their actual voices? What sounds like science fiction is now reality. Voice-Pro is an open-source Gradio WebUI that breaks the boundaries of audio manipulation. Powered by cutting-edge Whisper engines, this tool turns voice replication into child's play. Key Features: - Zero-shot Voice Cloning - Voice Changer with 50+ Celebrity Voices - YouTube Audio Downloading - Vocal Isolation - Multi-Language Text-to-Speech (Edge-TTS, F5-TTS) - Multi-Language Translation - Powered by Whisper Engines (Whisper, Faster-Whisper,…
2024 · github.com
- 15

- 16

- 17
- 18

- 19

- 20

- 21

- 22

- 23BC
I know there are various tools that are supposed to make this easy, but I couldn't find anything that did everything I wanted, so I made this today for fun. The web-based offerings all take forever and seem flaky, and you need to process one video at a time, with no control over the transcription settings. In contrast, my script lets you convert a whole playlist in bulk with full control over everything. It's truly easy to use-- you can clone the repo, install to a venv, and be generating a folder full of high quality transcript text files in under 5 minutes. All you need to do is supply the…
2023 · github.com
- 24

Ranked by how close each launch is in meaning, then by votes. Refine with a description →