nowfound

Alternatives

Products that do what Open-source turn detection model for voice AI does

Hey HN, it’s Russ - cofounder of LiveKit. An open source stack for building realtime AI applications. We’re sharing our first homegrown AI model for turn detection. Here’s a live demo: https://cerebras.vercel.app/ Voice AI has come a long way in the last year. We now have end-to-end systems that can generate a response to user input in 300-500ms — human level speeds! As latency reduces, a common problem that surfaces is the LLM responds too quickly. Any time there’s a short pause in a user’s speech, it ends up interrupting them. This is largely due to how voice AI applications…

  1. 1
    Ojin106

    Talk to an AI Agent with a real face and voice, in real time

    10d ago · ojin.ai

  2. 2

    Voice AI that feels as good as it sounds

    May 2026

  3. 3
    OpenWispr190

    100% local open source AI speech-to-text model

    2025

  4. 4
    Voxtral201

    Frontier open source speech understanding models

    2025

  5. 5RT

    Avaturn.live is a voice-to-voice AI assistant that lets you speak to an avatar in real time. We managed to reach <0.5 sec response time so that the conversation with an avatar feels more natural. Also what do you think about our lip sync? We were struggling a lot to reach our current level. Initial use case we're exploring: Automated product demos&#x2F;sales, where the AI avatar can showcase features and answer questions in real-time. Demo: https:&#x2F;&#x2F;dashboard.avaturn.live&#x2F;demo You can try talking to an avatar about anything

    2024 · dashboard.avaturn.live

  6. 6
    AI-Spy113

    AI audio detection

    2023

  7. 7

    The first multi-turn voice AI model built for conversation

    2024

  8. 8WO

    We kept hitting the same wall building voice AI systems. Pipecat and LiveKit are great projects, genuinely. But getting it to production took us weeks of plumbing - wiring things together, handling barge-ins, setting up telephony, Knowledge base, tool calls, handling barge in etc. And every time we needed to tweak agent behavior, you were back in the code and redeploying. We just wanted to change a prompt and test it in 30 seconds. Thats why Vapi retell etc exist. So we wrote the entire code and open sourced it as a Visual drag-and-drop for voice agents ( same as vapi or n8n for voice).…

    Mar 2026 · github.com

  9. 9OP

    Hey HN! Janak here from Outspeed (https:&#x2F;&#x2F;outspeed.com). We’re excited to show you Outspeed : a purpose-built platform for realtime voice & video AI applications. Here’s a demo of some cool apps you can create using Outspeed: https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=a11LQIlXelM Outspeed emerged from our frustration of needing to stitch together multiple tools such as livekit, vocode, langflow, silero etc. just to make a simple voice bot. Even after all that hard work, it still wasn’t production-ready. So we decided to work on a complete framework that could stand production…

    2024 · github.com

  10. 10CC

    Hey there HN! I believe the future of AI communication will be more voice and less text. Low-latency realistic voice interactions are finally becoming feasible. I've built a few voice-first apps on Retell AI using Elevenlabs voices. This one uses Claude Haiku for responses and Mixtral to switch between posts and comments. The AI knows about the top 30 posts and their comments on Hacker News right now. After a Google sign-in you can try it free for 10 minutes. I'd love to hear your thoughts!

    2024 · callhackernews.com

  11. 11BY

    Voice Activity Detection (VAD) is a crucial component for Voice AI, enabling more natural and efficient interactions. TEN VAD is an open-source solution designed to supercharge your Voice AI Agents with lightning-fast, human-like conversations! TEN VAD offers some key advantages: ONNX Support: Deploy on virtually any platform or hardware architecture! This means greater flexibility and easier integration into your existing systems. Superior Detection Accuracy: Experience noticeable improvements in voice detection, leading to fewer errors and more reliable performance. Smaller & Faster: Enjoy…

    2025 · github.com

  12. 12WH

    Hey folks. We built SAA (Selective Auditory Attention) after trying to find ways to make a good experience with multiple robots&#x2F;multiple agents. What typically ended up happening is they'd never stop talking. This is an SDK you can put before your STT. It lets you know when your device is being spoken to or not without a wakeword. You can use it for: -Single AI, Multi human -Multi AI, Single human -Multi AI, Multi human (we recommend also adding a wakeword on top for a better system) There are two models. One that is video + audio and one that is just audio. The way it overall works is…

    Jun 2026 · github.com

  13. 13BA

    cerebrium.ai&#x2F;blog&#x2F;how-to-build-a-real-time-ai-avatar-for-training-and-coaching At Cerebrium, we have recently built a few demos showing voice AI capabilities (worlds fastest voice agent & realtime RAG agent) but we wanted to push the boundary and see if we could create realistic, human-like situations to train and onboard teams to perform better - recreating real life scenarios! An example of this is a sales coach for your sales team, an investor pitch or even prep for a notoriously stressful YC interview . To achieve this there were a few difficult problems to solve, namely: - How…

    2024 · coaching.cerebrium.ai

  14. 14TA

    Hello HN, I am Brian Cardinale, a penetration tester and security researcher at SecureCoders. We have been performing more and more AI based security assessments. We were presented a unique challenge of testing a system where the only interface was voice based, and as much as I like talking on the phone , we decided to create a test harness to facilitate the actual testing in a more systematic way. The technical test harness was the easy part, though. Creating test goals and attack strategies to help facilitate repeated and comprehensive testing became the real challenge. As such, we have…

    Feb 2026 · redcaller.com

  15. 15OV

    Hi Hackernews, we're Maitreya, Prateek and Marmik. Over the past few months we've been working on building a platform to build, scale and monitor voice based LLM applications. Demo (https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=OSrOmyR7oQs) 1⃣ Open Source orchestration: We're open-sourcing our orchestration to quickly setup and create LLM based voice driven conversational applications https:&#x2F;&#x2F;github.com&#x2F;bolna-ai&#x2F;bolna&#x2F; 2⃣ Hosted API Platform: Exposing our managed solution via APIs to build voice driven applications…

    2024 · bolna.dev

  16. 16VG

    Hi, I'm Kamil and I'm a founder of Applied AI agency in Warsaw, Poland. We've trained a small <1MB voice classifier model that runs on CPU in 4ms. Can be run next to silero VAD in voice AI deployments. What we noticed in production deployments of voice assistants in Contact Centers in EU is that human consultants pick up immediately how to inflect verbs and ajdectives after one utterance from the caller. But voice AI agents don't know it until 1-2 minutes into the call when they are either corrected or the caller uses explicitly words with male&#x2F;female form a couple of times. Our model…

    May 2026 · huggingface.co

  17. 17RA

    Your Mac can run AI that holds its own against cloud models for the everyday stuff: chatting, making images, reading documents, transcribing voice. The hardware got there a while ago. The software to actually use it locally mostly didn't, so I built Off Grid. Download a model and it all runs on your machine. Ask it something on a flight with no wifi. Summarize a confidential document that never leaves your laptop. Run a hundred image generations in a loop and pay nothing, because it's your own GPU doing the work. Swap your paid dictation app for local Whisper. Talk through a coding problem…

    Jun 2026 · github.com

  18. 18ET

    For a while I've wanted to try out the new AI voices for long-form narration, but everything I found required a subscription that didn't justify my limited usage. I came across the open Kokoro model [0] and the voices are very good -- good enough to listen to for hours without the fatigue I got from legacy, robotic TTS voices. The model is 82m parameters and designed to run fast, but I still struggled to get reasonable times from CPU inference on my 12-core laptop. I thought a cloud-based GPU service would let me generate audiobooks fast enough to feed my own self-hosted library, and that…

    Jun 2026 · ebookaloud.com

  19. 19EE

    Hey everyone on HN! We recently spent the past couple of weeks building out an end-to-end platform which can plug-in multiple models (both open&#x2F;closed-source) to create voice driven conversational applications. We've tried to make the process simple & concise through documentation. Feel free to try it out and provide feedback. We will be launching a dashboard in the coming week for monitoring and analytics alongwith more open source models. Let us know what you all think. (if you want to contribute, we have tons of features planned - do let us know)

    2023 · github.com

  20. 20VG

    Hey HN, I’m Asa, one of the co-founders of Voicepanel (https:&#x2F;&#x2F;voicepanel.co). We’re building a new way to gather feedback using AI, as an alternative to running surveys. Simply share what you want feedback on; Voicepanel will recruit respondents, conduct open-ended interviews, and synthesize the insights automatically. Check out a demo here: https:&#x2F;&#x2F;youtu.be&#x2F;6_C8netV38E, sign up for a free 7 day trial (https:&#x2F;&#x2F;voicepanel.co&#x2F;) or try out examples of our product in action without a login (https:&#x2F;&#x2F;www.voicepanel.co&#x2F;#examples). Why…

    2024

  21. 21WH
  22. 22AA

    Hi HN, I'm the creator of AVA - AI Voice Agent for Asterisk My repo was shared here once before by someone else so I wanted to follow up with the progress since then. https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46380399 I've been working with Asterisk&#x2F;FreePBX systems for years. I wanted to add AI voice capabilities to legacy phone systems without paying per-minute SaaS fees or ripping out the entire telephony stack. So I built AVA, a self-hosted AI voice agent that can integrate into any traditional phone system. While most solutions demand expensive migrations to cloud-only…

    Mar 2026 · github.com

  23. 23

    Build voice agents you own

    Jul 2026 · dashboard.voice-ai.dev

  24. 24AE

    Hey HN! Have been working on Audentic, a platform that simplifies adding voice AI to your website. Think of it as a copy-paste voice assistant that you can embed directly into your site. Setting up voice AI has traditionally been a bit of a hassle, often requiring juggling multiple components like speech recognition, text processing, and text-to-speech systems. With Audentic, we've streamlined this into a more straightforward, end-to-end voice model. While OpenAI's Realtime API has made real-time, multimodal AI interactions more feasible, integrating these capabilities into a website can…

    2025 · audentic.io

Ranked by how close each launch is in meaning, then by votes. Refine with a description →