Alternatives
Products that do what Whisper v3 API does
I'm working on making it easier to use open-source AI models by providing simple APIs. Last week, OpenAI released Whisper v3, which showed improved performance across all its 100+ supported languages. We tried to optimize it as much as possible to be able to offer it for $0.0028 per minute (vs $0.006 on OpenAI). I hope it's helpful for someone: https://www.lemonfox.ai/apis/speech-to-text
- 1

- 2
- 301
Hey HN! I've been working on a side project to create an audio transcription API based on the OpenAI whisper model. Sign up link: https://whisperapi.com I tried to make the API really easy to use and get setup with. Also, because the Whisper model is so good, turns out I can offer the service for about 75% cheaper than what seems like the industry average. I'm always looking to make improvements, so would appreciate any feedback anyone has!
2022 · whisperapi.com
- 4

- 5

- 6

- 7

- 8

- 9

- 10

- 11
- 12

- 13

- 14

- 15OS
I built Whispering because I believe transcription is too fundamental a tool to be locked behind paywalls. It's a cross-platform desktop and web transcription app that turns speech into text with a keyboard shortcut, among other things. The app lets you bring your own API key (OpenAI, Groq, etc.) and make direct calls. If you want complete privacy, it also supports local transcription. Either way, your audio never goes through any middleman servers. It's super lightweight (~22MB), built with Svelte 5 and Tauri, and works on Mac, Windows, and Linux. I've been using it daily for the past few…
2025 · github.com
- 16

- 17

- 18ES
Are you spending hundreds of dollars a month on AI coding costs? I built European Swallow AI, an API that uses reasoning models (Claude, Deepseek) for thinking and cheaper specialized coding models (Qwen, Grok) to write code, so you can save token costs while still getting high quality code. With an OpenAI formatted endpoint you can try European Swallow in Cursor, Typing Mind, Xibe AI and your own custom apps. During testing, European Swallow scored 80.5% on Big Code Bench and over 90% on the HumanEval+. It averaged $2.60 per million tokens compared with the $15 per million output tokens of…
Oct 2025 · europeanswallowai.com
- 19

- 20VF
Hey HN, I'm Josiah. We love voice dictation, but wanted an open source version for transparency, privacy, and something that everyone could contribute to. So we built Voquill, an open source alternative to WisprFlow, Monologue, and Willow. It lets you dictate into any desktop app. Press a hotkey, talk, text gets inserted. You can run Whisper locally, use our server, or wire up any provider you want (OpenAI, Claude, Groq, OpenRouter, whatever). You have full control over where your data goes. Runs on Windows, macOS, and Linux. Open source, AGPLv3, built with Tauri and Rust. We're working on a…
Feb 2026 · github.com
- 21OW
I built Open WhisperScribe after struggling with the complexity, paywalls, or limitations of most speech-to-text tools for macOS and elsewhere. My main requirements were: Simple, <5-min setup Runs fully offline, locally (no data leaves your machine) Works seamlessly in any app — just press a hotkey, speak, and your words appear where your cursor is (code editors, terminal, you name it) Fast and distraction-free It uses OpenAI’s Whisper model under the hood, but wraps it in a lightweight CLI tool that sits quietly in the background. The project is open source (Apache 2.0). Setup is a single…
2025 · github.com
- 22CM
Hi HN, Today, we're launching Glowby Basic, our open-source, customizable AI assistant powered by OpenAI API. It allows you to easily recreate and adapt your own assistant for various needs, with multilingual voice capabilities. Key features include autonomous decision-making, customizable pre-set questions, support for 30+ languages, and effortless prompt switching for different tasks and scenarios. We've made interactions with LLMs and AI-related APIs open-source to encourage collaboration and enhance transparency. Try our live demo (no sign-up required) and explore the GitHub repo linked…
2023 · github.com
- 23AA
The audio/TTS space just moved fast. In the last week alone: NVIDIA – PersonaPlex-7B Open-source, full-duplex conversational speech model. Inworld AI – TTS-1.5 Realtime TTS (<250ms), $0.005/min, currently #1 on Artificial Analysis. Flash Labs – Chroma 1.0 First open-source, end-to-end, real-time speech-to-speech model. Alibaba Qwen – Qwen3-TTS Fully open-sourced TTS family: Base, CustomVoice, VoiceDesign. Kyutai Labs – Pocket TTS Runs locally on a laptop. No GPU required. Feels like TTS is hitting the same acceleration moment LLMs had last year. Realtime, open-source, and local is…
Jan 2026 · github.com
- 24WO
We kept hitting the same wall building voice AI systems. Pipecat and LiveKit are great projects, genuinely. But getting it to production took us weeks of plumbing - wiring things together, handling barge-ins, setting up telephony, Knowledge base, tool calls, handling barge in etc. And every time we needed to tweak agent behavior, you were back in the code and redeploying. We just wanted to change a prompt and test it in 30 seconds. Thats why Vapi retell etc exist. So we wrote the entire code and open sourced it as a Visual drag-and-drop for voice agents ( same as vapi or n8n for voice).…
Mar 2026 · github.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →