Alternatives
Products that do what I open-sourced my AI toy company that runs on ESP32 and OpenAI realtime does
Hi HN! Last year the project I launched here got a lot of good feedback on creating speech to speech AI on the ESP32. Recently I revamped the whole stack, iterated on that feedback and made our project fully open-source—all of the client, hardware, firmware code. This Github repo turns an ESP32-S3 into a realtime AI speech companion using the OpenAI Realtime API, Arduino WebSockets, Deno Edge Functions, and a full-stack web interface. You can talk to your own custom AI character, and it responds instantly. I couldn't find a resource that helped set up a reliable, secure websocket (WSS) AI…
- 1

- 2

Generating uncanny AI avatars is now open source
May 2026 · avaturn.live
- 3RT
Related: https://news.ycombinator.com/item?id=47653752
Apr 2026 · github.com
- 4

- 5

- 6
- 7

- 8

- 9

- 10

- 11

- 12

- 13

- 14

The most accurate streaming speech model for voice agents.
Mar 2026
- 15

- 16

- 17WV
I'm working on making it easier to use open-source AI models by providing simple APIs. Last week, OpenAI released Whisper v3, which showed improved performance across all its 100+ supported languages. We tried to optimize it as much as possible to be able to offer it for $0.0028 per minute (vs $0.006 on OpenAI). I hope it's helpful for someone: https://www.lemonfox.ai/apis/speech-to-text
2023
- 18SH
Hi HN folks ! I am the author of AVA, a self hosted AI Voice Agent that plugs into Asterisk/Freepbx so you own all the aspects of an AI Voice agent in your own infrastructure. It uses Asterisk native Audiosocket/RTP with python engine to run STT,LLM and TTS loop. The project support several full providers openai, gemini, grok, elevenlabs out of the box and also provides options to build custom pipelines by choosing different stt tts and llm. It also supports full local agent if you have a GPU with 25GB RAM which enables realtime conversation along with tool calling. I started this…
Jul 2026 · github.com
- 19WH
6d ago · fellowgeek.github.io
- 20

- 21PG
we have been building an open source orchestration which enables you to plug in your own TTS/ASR/LLM for end-to-end voice conversations at -> https://github.com/bolna-ai/bolna. Few days back, was having a discussion here in HN about the possibilities of having a complete open source stack for ASR+LLM+TTS. Today, we are releasing a complete open sourced Dockerized stack by merging Bolna with Whisper ASR, Llama3 and Melo TTS.
2024 · github.com
- 22ET
For a while I've wanted to try out the new AI voices for long-form narration, but everything I found required a subscription that didn't justify my limited usage. I came across the open Kokoro model [0] and the voices are very good -- good enough to listen to for hours without the fatigue I got from legacy, robotic TTS voices. The model is 82m parameters and designed to run fast, but I still struggled to get reasonable times from CPU inference on my 12-core laptop. I thought a cloud-based GPU service would let me generate audiobooks fast enough to feed my own self-hosted library, and that…
Jun 2026 · ebookaloud.com
- 23AA
The audio/TTS space just moved fast. In the last week alone: NVIDIA – PersonaPlex-7B Open-source, full-duplex conversational speech model. Inworld AI – TTS-1.5 Realtime TTS (<250ms), $0.005/min, currently #1 on Artificial Analysis. Flash Labs – Chroma 1.0 First open-source, end-to-end, real-time speech-to-speech model. Alibaba Qwen – Qwen3-TTS Fully open-sourced TTS family: Base, CustomVoice, VoiceDesign. Kyutai Labs – Pocket TTS Runs locally on a laptop. No GPU required. Feels like TTS is hitting the same acceleration moment LLMs had last year. Realtime, open-source, and local is…
Jan 2026 · github.com
- 24

Ranked by how close each launch is in meaning, then by votes. Refine with a description →