nowfound

Alternatives

Products that do what Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B does

Related: https://news.ycombinator.com/item?id=47653752

  1. 1
    Play AI447

    The voice interface of AI

    2024

  2. 2

    Premium AI voice quality without the premium price tag.

    2025

  3. 3IO

    Hi HN! Last year the project I launched here got a lot of good feedback on creating speech to speech AI on the ESP32. Recently I revamped the whole stack, iterated on that feedback and made our project fully open-source—all of the client, hardware, firmware code. This Github repo turns an ESP32-S3 into a realtime AI speech companion using the OpenAI Realtime API, Arduino WebSockets, Deno Edge Functions, and a full-stack web interface. You can talk to your own custom AI character, and it responds instantly. I couldn't find a resource that helped set up a reliable, secure websocket (WSS) AI…

    2025 · github.com

  4. 4

    For reliable, production-ready voice agents

    2025

  5. 5RT
  6. 6

    Create realistic AI Voiceovers within seconds

    2022

  7. 7
    Wavel AI289

    Most realistic and natural AI voiceover & dubbing for videos

    2023

  8. 8

    Free text-to-speech generator with realistic AI voices

    2024

  9. 9

    Quickly turn any song, voice memo or podcast into a video

    2025

  10. 10

    Use your voice to create realistic AI speech in real-time

    2022

  11. 11OS

    Heeey! I built a macOS copilot that has been useful to me, so I open sourced it in case others would find it useful too. It's pretty simple: - Use a keyboard shortcut to take a screenshot of your active macOS window and start recording the microphone. - Speak your question, then press the keyboard shortcut again to send your question + screenshot off to OpenAI Vision - The Vision response is presented in-context/overlayed over the active window, and spoken to you as audio. - The app keeps running in the background, only taking a screenshot/listening when activated by keyboard…

    2023 · github.com

  12. 12

    The voice AI auto-testing loop to simulate, evaluate & ship

    2025

  13. 13

    Run multimodal AI locally with an encoder-free architecture

    Jun 2026 · blog.google

  14. 14

    High-quality voice clones with just 60 seconds of audio

    2023

  15. 15

    Voice AI that feels as good as it sounds

    May 2026 · inworld.ai

  16. 16
    Audino AI325

    Make content creation simpler with AI-generated audio

    2025

  17. 17
    Allinpod264

    The future of AI podcasting

    2023

  18. 18
    Gemma 3n199

    Run powerful multimodal AI right on your phone

    2025

  19. 19

    Voice agents powered by Simba 3.2 the world's #1 voice model

    Jul 2026 · speechify.ai

  20. 20

    Real-time audio conversations on-device

    Oct 2025

  21. 21G4

    About six months ago, I started working on a project to fine-tune Whisper locally on my M2 Ultra Mac Studio with a limited compute budget. I got into it. The problem I had at the time was I had 15,000 hours of audio data in Google Cloud Storage, and there was no way I could fit all the audio onto my local machine, so I built a system to stream data from my GCS to my machine during training. Gemma 3n came out, so I added that. Kinda went nuts, tbh. Then I put it on the shelf. When Gemma 4 came out a few days ago, I dusted it off, cleaned it up, broke out the Gemma part from the Whisper…

    Apr 2026 · github.com

  22. 22
    AI-Spy113

    AI audio detection

    2023

  23. 23IT
  24. 24MP

    I work on real-time voice/video AI at Tavus and for the past few years, I’ve mostly focused on how machines respond in a conversation. One thing that’s always bothered me is that almost all conversational systems still reduce everything to transcripts, and throw away a ton of signals that need to be used downstream. Some existing emotion understanding models try to analyze and classify into small sets of arbitrary boxes, but they either aren’t fast / rich enough to do this with conviction in real-time. So I built a multimodal perception system which gives us a way to encode visual…

    Feb 2026 · raven.tavuslabs.org

Ranked by how close each launch is in meaning, then by votes. Refine with a description →