Alternatives
Products that do what Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B does
Related: https://news.ycombinator.com/item?id=47653752
- 1

- 2

- 3IO
Hi HN! Last year the project I launched here got a lot of good feedback on creating speech to speech AI on the ESP32. Recently I revamped the whole stack, iterated on that feedback and made our project fully open-source—all of the client, hardware, firmware code. This Github repo turns an ESP32-S3 into a realtime AI speech companion using the OpenAI Realtime API, Arduino WebSockets, Deno Edge Functions, and a full-stack web interface. You can talk to your own custom AI character, and it responds instantly. I couldn't find a resource that helped set up a reliable, secure websocket (WSS) AI…
2025 · github.com
- 4

- 5RT
2025 · github.com
- 6

- 7

- 8

- 9

- 10

- 11OS
Heeey! I built a macOS copilot that has been useful to me, so I open sourced it in case others would find it useful too. It's pretty simple: - Use a keyboard shortcut to take a screenshot of your active macOS window and start recording the microphone. - Speak your question, then press the keyboard shortcut again to send your question + screenshot off to OpenAI Vision - The Vision response is presented in-context/overlayed over the active window, and spoken to you as audio. - The app keeps running in the background, only taking a screenshot/listening when activated by keyboard…
2023 · github.com
- 12

- 13

Run multimodal AI locally with an encoder-free architecture
Jun 2026 · blog.google
- 14

- 15

- 16

- 17

- 18

- 19

Voice agents powered by Simba 3.2 the world's #1 voice model
Jul 2026 · speechify.ai
- 20

- 21G4
About six months ago, I started working on a project to fine-tune Whisper locally on my M2 Ultra Mac Studio with a limited compute budget. I got into it. The problem I had at the time was I had 15,000 hours of audio data in Google Cloud Storage, and there was no way I could fit all the audio onto my local machine, so I built a system to stream data from my GCS to my machine during training. Gemma 3n came out, so I added that. Kinda went nuts, tbh. Then I put it on the shelf. When Gemma 4 came out a few days ago, I dusted it off, cleaned it up, broke out the Gemma part from the Whisper…
Apr 2026 · github.com
- 22

- 23IT
2025 · github.com
- 24MP
I work on real-time voice/video AI at Tavus and for the past few years, I’ve mostly focused on how machines respond in a conversation. One thing that’s always bothered me is that almost all conversational systems still reduce everything to transcripts, and throw away a ton of signals that need to be used downstream. Some existing emotion understanding models try to analyze and classify into small sets of arbitrary boxes, but they either aren’t fast / rich enough to do this with conviction in real-time. So I built a multimodal perception system which gives us a way to encode visual…
Feb 2026 · raven.tavuslabs.org
Ranked by how close each launch is in meaning, then by votes. Refine with a description →