SpeakoFlow Mini
Tiny local model that fixes dictation without rewriting it
What it does
Dictation cleanup has two halves. Fix what the speaker actually got wrong, and leave everything else exactly as dictated. General models fail the second half: hand them a sentence that is already correct and they improve it. A comma becomes a full stop, a paragraph becomes bullets. SpeakoFlow Mini is 0.8B, Apache-2.0, 833 MB, and runs offline on a desktop CPU. Already-correct text comes back untouched 92.6% of the time. No API, no reasoning tokens, no rewriting.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Dictation cleanup for transcribed speech. It applies the correction the speaker actually made, and leaves everything else exactly as you said it. 833 MB at Q8_0, 2,509 ms median on a desktop CPU. Fine-tuned from Qwen/Qwen3.5-0.8B with LoRA rank 16, merged, then quantised. English. Not a chat model, not a rewriter. Ships in SpeakoFlow , a free offline voice assistant for Windows, macOS and Linux. Cleanup splits cleanly into work a rule can do and work it cannot. Stage one is deterministic. Filler words, repeated words, spacing, punctuation, capitalisation, numbers, dates, currency and known jargon substitutions are pattern work, and pattern work belongs in code, where it is fast, free and…from huggingface.co
Does the same job
all alternatives →
SpeakoFlowAug 2026 · speakoflow.com · ▲95Free, open-source voice dictation and AI assistant, offline



More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 16d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 18d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 26d ago · cactuscompute.com


Launched alongside, August 2026
the whole month →- TL
Life & fun · 10d ago · louisabraham.github.io


- SA
Hello HN! I found that picking out plausible but diverse skin tones for my digital art and game development projects was kind of difficult, and I got curious about if there was a way to define a color space that made it easy. I've built a color picker and procedural generation algorithm based on the space as well as a bunch of other fun js features and demos throughout the page that use the equations. If you find it interesting, I have lots of explanations of how I built it and what properties the space has. The methodology might be a bit shaky, but hopefully the result is as helpful for…
Life & fun · Aug 2026 · toneyalexander.github.io


I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 16d ago · simedw.com