Alternatives
Products that do what Audio AI had a wild day – 5 major open-source / real-time TTS drops does
The audio/TTS space just moved fast. In the last week alone: NVIDIA – PersonaPlex-7B Open-source, full-duplex conversational speech model. Inworld AI – TTS-1.5 Realtime TTS (<250ms), $0.005/min, currently #1 on Artificial Analysis. Flash Labs – Chroma 1.0 First open-source, end-to-end, real-time speech-to-speech model. Alibaba Qwen – Qwen3-TTS Fully open-sourced TTS family: Base, CustomVoice, VoiceDesign. Kyutai Labs – Pocket TTS Runs locally on a laptop. No GPU required. Feels like TTS is hitting the same acceleration moment LLMs had last year. Realtime, open-source, and local is…
- 1KT
Kitten TTS is an open-source series of tiny and expressive text-to-speech models for on-device applications. We are excited to launch a preview of our smallest model, which is less than 25 MB. This model has 15M parameters. This release supports English text-to-speech applications in eight voices: four male and four female. The model is quantized to int8 + fp16, and it uses onnx for runtime. The model is designed to run literally anywhere eg. raspberry pi, low-end smartphones, wearables, browsers etc. No GPU required! We're releasing this to give early users a sense of the latency and voices…
2025 · github.com
- 2

- 3
- 4

- 5TN
Kitten TTS (https://github.com/KittenML/KittenTTS) is an open-source series of tiny and expressive text-to-speech models for on-device applications. We had a thread last year here: https://news.ycombinator.com/item?id=44807868. Today we're releasing three new models with 80M, 40M and 14M parameters. The largest model (80M) has the highest quality. The 14M variant reaches new SOTA in expressivity among similar sized models, despite being <25MB in size. This release is a major upgrade from the previous one and supports English text-to-speech applications in…
Mar 2026 · github.com
- 6

- 7

- 8

- 9RT
2025 · github.com
- 10

First TTS model to support all 22 Indic languages + English
2024
- 11

- 12MO
I wanted to share our new speech to text model, and the library to use them effectively. We're a small startup (six people, sub-$100k monthly GPU budget) so I'm proud of the work the team has done to create streaming STT models with lower word-error rates than OpenAI's largest Whisper model. Admittedly Large v3 is a couple of years old, but we're near the top the HF OpenASR leaderboard, even up against Nvidia's Parakeet family. Anyway, I'd love to get feedback on the models and software, and hear about what people might build with it.
Feb 2026 · github.com
- 13

- 14

- 15

- 16IO
Hi HN! Last year the project I launched here got a lot of good feedback on creating speech to speech AI on the ESP32. Recently I revamped the whole stack, iterated on that feedback and made our project fully open-source—all of the client, hardware, firmware code. This Github repo turns an ESP32-S3 into a realtime AI speech companion using the OpenAI Realtime API, Arduino WebSockets, Deno Edge Functions, and a full-stack web interface. You can talk to your own custom AI character, and it responds instantly. I couldn't find a resource that helped set up a reliable, secure websocket (WSS) AI…
2025 · github.com
- 17

- 18

- 19

- 20

- 21

- 22TA
2012 · tts-api.com
- 23

- 24

Ranked by how close each launch is in meaning, then by votes. Refine with a description →