nowfound

Alternatives

Products that do what Agent Wispr does

Local Whisper dictation built for coding terminals

  1. 1
    Wispr Flow2,737

    Speak naturally, write perfectly & 3x faster in every app

    2024 · wisprflow.ai

  2. 2
    OpenWispr190

    100% local open source AI speech-to-text model

    2025

  3. 3FA
  4. 4KT

    Kitten TTS is an open-source series of tiny and expressive text-to-speech models for on-device applications. We are excited to launch a preview of our smallest model, which is less than 25 MB. This model has 15M parameters. This release supports English text-to-speech applications in eight voices: four male and four female. The model is quantized to int8 + fp16, and it uses onnx for runtime. The model is designed to run literally anywhere eg. raspberry pi, low-end smartphones, wearables, browsers etc. No GPU required! We're releasing this to give early users a sense of the latency and voices…

    2025 · github.com

  5. 5PO

    Hi HN, OpenAI recently released a model for automatic speech recognition called Whisper [0]. I decided to reimplement the inference of the model from scratch using C/C++. To achieve this I implemented a minimalistic tensor library in C and ported the high-level architecture of the model in C++. The entire code is less than 8000 lines of code and is contained in just 2 source files without any third-party dependencies. The Github project is here: https://github.com/ggerganov/whisper.cpp With this implementation I can very easily build and run the model - “make…

    2022 · github.com

  6. 6

    Local AI dictation for Windows

    19d ago · whisperstream.io

  7. 7

    Build Powerful Voice Agents

    2025

  8. 8

    AI dictation that turns messy speech into polished text.

    Feb 2026

  9. 9
    Lispr240

    Hold a key, speak, and Lispr writes it anywhere

    Jul 2026 · lispr.ai

  10. 10IB

    I built a voice agent from scratch that averages ~400ms end-to-end latency (phone stop → first syllable). That’s with full STT → LLM → TTS in the loop, clean barge-ins, and no precomputed responses. What moved the needle: Voice is a turn-taking problem, not a transcription problem. VAD alone fails; you need semantic end-of-turn detection. The system reduces to one loop: speaking vs listening. The two transitions - cancel instantly on barge-in, respond instantly on end-of-turn - define the experience. STT → LLM → TTS must stream. Sequential pipelines are dead on arrival for natural…

    Mar 2026 · ntik.me

  11. 11
    Zro428

    Private inference for coding agents

    Jul 2026 · zro.moonmath.ai

  12. 12MO

    I wanted to share our new speech to text model, and the library to use them effectively. We're a small startup (six people, sub-$100k monthly GPU budget) so I'm proud of the work the team has done to create streaming STT models with lower word-error rates than OpenAI's largest Whisper model. Admittedly Large v3 is a couple of years old, but we're near the top the HF OpenASR leaderboard, even up against Nvidia's Parakeet family. Anyway, I'd love to get feedback on the models and software, and hear about what people might build with it.

    Feb 2026 · github.com

  13. 13
    Wispli67

    Speed of Voice. Power of AI.

    Apr 2026 · wispli.com

  14. 14

    Local transcripts with speaker labels, timestamps, + export

    Dec 2025

  15. 15AT

    A 3.16M-parameter INT4 transformer running entirely in the on-chip memory of a Xilinx Kria KV260. Zero DRAM in the token loop, 59,965 tok/s on the fabric, bit-exact. Chat with it live.

    27d ago · mikeayles.com

  16. 16IM
  17. 17

    A neural net for speech recognition

    2022

  18. 18
    Wispro143

    Stop typing, start talking, get perfectly written text

    Jul 2026 · wisproapp.com

  19. 19OS

    Our goal with this project is to build a completely open source, state of the art turn detection model that can be used in any voice AI application. I've been experimenting with LLM voice conversations since GPT-4 was first released. (There's a previous front page Show HN about Pipecat, the open source voice AI orchestration framework I work on. [1]) It's been almost two years, and for most of that time, I've been expecting that someone would "solve" turn detection. We all built initial, pretty good 80/20 versions of turn detection on top of VAD (voice activity detection) models. And…

    2025 · github.com

  20. 20EL
  21. 21

    The fastest generative AI Text-to-Speech API

    2023

  22. 22

    Talk to any app on your Mac

    Jul 2026 · wisprkey.com

  23. 23WP

    This project is a Windows port of the whisper.cpp implementation: https://github.com/ggerganov/whisper.cpp Which in turn is a C++ port of OpenAI's Whisper automatic speech recognition (ASR) model: https://github.com/openai/whisper The implementation has no dependencies, usually much faster than realtime, and should hopefully work on most Windows computers in the world.

    2023 · github.com

  24. 24

    Production ASR for noisy multilingual audio

    Apr 2026 · microsoft.ai

Ranked by how close each launch is in meaning, then by votes. Refine with a description →