nowfound

Alternatives

Products that do what StreamKit does

Build and run Live video, speech-to-text, voice agent,

  1. 1

    Build massive-scale, real-time video and audio experiences

    2022

  2. 2
    LiveKit236

    The open source platform for real-time communication

    2021

  3. 3IB

    I built a voice agent from scratch that averages ~400ms end-to-end latency (phone stop → first syllable). That’s with full STT → LLM → TTS in the loop, clean barge-ins, and no precomputed responses. What moved the needle: Voice is a turn-taking problem, not a transcription problem. VAD alone fails; you need semantic end-of-turn detection. The system reduces to one loop: speaking vs listening. The two transitions - cancel instantly on barge-in, respond instantly on end-of-turn - define the experience. STT → LLM → TTS must stream. Sequential pipelines are dead on arrival for natural…

    Mar 2026 · ntik.me

  4. 4

    Build Powerful Voice Agents

    2025

  5. 5VP

    Imagine creating a podcast where Mark Zuckerberg interviews Elon Musk – using their actual voices? What sounds like science fiction is now reality. Voice-Pro is an open-source Gradio WebUI that breaks the boundaries of audio manipulation. Powered by cutting-edge Whisper engines, this tool turns voice replication into child's play. Key Features: - Zero-shot Voice Cloning - Voice Changer with 50+ Celebrity Voices - YouTube Audio Downloading - Vocal Isolation - Multi-Language Text-to-Speech (Edge-TTS, F5-TTS) - Multi-Language Translation - Powered by Whisper Engines (Whisper, Faster-Whisper,…

    2024 · github.com

  6. 6

    Next gen audio/video editor with personalized voice cloning

    2020

  7. 7

    Real-time text-to-speech model you can self-host

    May 2026 · kugelaudio.com

  8. 8OS

    Hey HN, we've been working with OpenAI for the past few months on the new Realtime API. The goal is to give everyone access to the same stack that underpins Advanced Voice in the ChatGPT app. Under the hood it works like this: - A user's speech is captured by a LiveKit client SDK in the ChatGPT app - Their speech is streamed using WebRTC to OpenAI’s voice agent - The agent relays the speech prompt over websocket to GPT-4o - GPT-4o runs inference and streams speech packets (over websocket) back to the agent - The agent relays generated speech using WebRTC back to the user’s device The…

    2024 · github.com

  9. 9LV
  10. 10

    Voice AI that feels as good as it sounds

    May 2026 · inworld.ai

  11. 11

    Voice agents powered by Simba 3.2 the world's #1 voice model

    Jul 2026 · speechify.ai

  12. 12

    The most accurate streaming speech model for voice agents.

    Mar 2026

  13. 13

    Generate endless looping sound effects with a single prompt

    2025

  14. 14IM
  15. 15

    Dub & translate any video in any language

    2024

  16. 16

    Record, capture and transcribe audio & video in minutes

    2021

  17. 17

    Convert audio and video to accurate text in seconds with AI

    2024

  18. 18

    One API to build production-ready voice agents

    Apr 2026 · assemblyai.com

  19. 19
    Aloud118

    Turn spoken feedback into tasks your coding agent can run

    17d ago · aloud.sh

  20. 20

    Fast, accurate STT and TTS APIs at the best price

    Apr 2026 · x.ai

  21. 21
    Speechius127

    The teleprompter that actually listens

    Jul 2026 · speechius.com

  22. 22

    Accurate transcription and translation for video pros

    2018

  23. 23

    Deploy your LiveKit Voice AI agents instantly

    2025

  24. 24

    Open-Source Multilingual Text-to-Speech with Voice Cloning

    2024

Ranked by how close each launch is in meaning, then by votes. Refine with a description →