nowfound

Alternatives

Products that do what I trained a 9M speech model to fix my Mandarin tones does

Built this because tones are killing my spoken Mandarin and I can't reliably hear my own mistakes. It's a 9M Conformer-CTC model trained on ~300h (AISHELL + Primewords), quantized to INT8 (11 MB), runs 100% in-browser via ONNX Runtime Web. Grades per-syllable pronunciation + tones with Viterbi forced alignment. Try it here: https://simedw.com/projects/ear/

  1. 1MT
  2. 2
    yuyin.io136

    Master Chinese pronunciation with AI-powered feedback

    2025

  3. 3IB

    Built a ~9M param LLM from scratch to understand how they actually work. Vanilla transformer, 60K synthetic conversations, ~130 lines of PyTorch. Trains in 5 min on a free Colab T4. The fish thinks the meaning of life is food. Fork it and swap the personality for your own character.

    Apr 2026 · github.com

  4. 4

    Multilingual speech AI model trained on 12.5M hours of data

    2024

  5. 5IB

    I built a voice agent from scratch that averages ~400ms end-to-end latency (phone stop → first syllable). That’s with full STT → LLM → TTS in the loop, clean barge-ins, and no precomputed responses. What moved the needle: Voice is a turn-taking problem, not a transcription problem. VAD alone fails; you need semantic end-of-turn detection. The system reduces to one loop: speaking vs listening. The two transitions - cancel instantly on barge-in, respond instantly on end-of-turn - define the experience. STT → LLM → TTS must stream. Sequential pipelines are dead on arrival for natural…

    Mar 2026 · ntik.me

  6. 6

    Fast, accurate STT for production-grade voice agents

    May 2026 · ringg.ai

  7. 7

    First TTS model to support all 22 Indic languages + English

    2024

  8. 8DS

    I've been working on a little side project that combines Duolingo-like listening comprehension exercises with real content . Every video is transcribed to get much better transcripts than the closed captions. I filter on high quality transcripts, and afterwards a LLM selects only plausible segments for the exercises. This seems to work well for quality control and seems to be reliable enough for these short exercises. Would love your thoughts!

    2025 · app.fluentsubs.com

  9. 9

    Master Chinese pronunciation with real-time AI feedback

    Feb 2026

  10. 101M
  11. 11
    Syllabics100

    Improve your English pronunciation

    2022

  12. 12

    A native omni model for voice, video, and tools

    Mar 2026

  13. 13SM

    Hey HN! We built Speech Meter as a tool to practice and improve English pronunciation. It uses AI to analyze your accent and score your pronunciation accuracy. It’s great for anyone who wants to practice their pronunciation in English. I’d love to hear your thoughts and suggestions for improvements. I really appreciate any feedback you could have (:

    2023 · speechmeter.com

  14. 14
    YiChi94

    Learn to speak fluent Chinese Mandarin

    2021

  15. 15IB

    Hello Hacker News, When learning foreign languages, I made the most progress by speaking them throughout the day, every day. So I made a site where you can *speak* to an AI language teacher to practice both listening and speaking. # The product *What I have now:* * Multilingual speech recognition: You can ask a question in English and get an answer in your target language. * Feedback on your grammar. * Suggestions: See examples of what to say next to keep the conversation flowing. * Speed: Choose a lower speed for beginners or a faster one for advanced levels. * Translations: Click to see a…

    2023 · gliglish.com

  16. 16TN

    Here is a tool I built initially for myself to help with my German and Greek language studies. It started as a hack for creating Anki cards from native language audio. It extracts the words, finds their base forms (lemmas) and groups the examples by the lemma. At some point I realised that I have a transcription with word level timestamps that opens a lot of other opportunities. So I added a mode to click the first and last word in the transcript and it starts looping with the right gap and repeat count. Another feature I use a lot is selecting an audio fragment, sending a predefined prompt…

    Jun 2026 · lingochunk.com

  17. 17MW

    I've built mandoBot, a web app that segments and translates Mandarin Chinese text. This is a Django API (using Django-Ninja and PostgreSQL) and a NextJS front-end (with Typescript and Chakra). For a sample of what this app does, head to https://mandobot.netlify.app/?share_id=e8PZ8KFE5Y. This is my presentation of the first chapter of a classic story from the Republican era of Chinese fiction, Diary of a Madman by Lu Xun. Other chapters are located in the "Reading Room" section of the app. This app exists because reading Mandarin is very hard for learners (like me), since…

    2025 · mandobot.netlify.app

  18. 18FG

    We developed a new framework that enables flexible control of generated text in language models. By combining several models and/or system prompts in one mathematical formula, it lets you tweak your style and combine model outputs with ease. A handy tool for those working with LLMs, looking for more fine-grained control of stylistic output. More details in our paper: https://arxiv.org/abs/2311.14479. Feedback and potential applications are welcome.

    2023 · github.com

  19. 19

    Production ASR for noisy multilingual audio

    Apr 2026

  20. 20LL

    Hi! I'm living as an expat in France and I had a hard time following the news. So I've been working on Fluentsubs that turns YouTube content into a language learning experience. YouTube videos up to 20 minutes can be transcribed and listened to. At the same time I made it easy to lookup certain words and automatically add spaced repetition cards using the given context. This means that you can also rehearse with a real voice and not with AI dubs. This works especially well for the somewhat smaller languages that have a small track on Duolingo, and now are desperately looking for content…

    2025 · fluentsubs.com

  21. 21OS

    Our goal with this project is to build a completely open source, state of the art turn detection model that can be used in any voice AI application. I've been experimenting with LLM voice conversations since GPT-4 was first released. (There's a previous front page Show HN about Pipecat, the open source voice AI orchestration framework I work on. [1]) It's been almost two years, and for most of that time, I've been expecting that someone would "solve" turn detection. We all built initial, pretty good 80/20 versions of turn detection on top of VAD (voice activity detection) models. And…

    2025 · github.com

  22. 22AA

    2018 · play.google.com

  23. 23MA

    Hi HN! In trying to improve my Chinese skills, I built this tool that lets me write a mix of Chinese and English, then recommends a proper Chinese expression. This is super helpful when I want to write something in Chinese but I don't know all the vocab/grammar - I can enter my best effort and use English for the parts I don't know. Really, the fundamental benefit of the tool is that it encourages me to exercise the writing muscle, rather than defaulting to translating from English. My goal was to build something that is fast, relatively inexpensive, and not prone to misleading people.…

    2024 · unscrambler.dpw.me

  24. 24

    Physical mouth retraining, not more audio drills

    Jul 2026 · etsy.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →