Alternatives
Products that do what I fine-tuned Qwen 3.5 (0.8B–4B) on a Mac for text-to-SQL – 2B beats 12B does
- 1
Qwen3.5 Small▲3590.8B-9B native multimodal w/ more intelligence, less compute
Mar 2026 · huggingface.co
- 2

I built slotstream, a way to run Qwen3.8-Flash-Next 4-bit on a low-memory mac starting from 16GB, a 125B parameter model that would need 100GB+ memory/RAM, thanks to expert-offloading/ssd-streaming. Easy to install/update, and mac-native using MLX and Swift. It ships with auto-mode, which makes a good tradeoff between memory usage and speed. I'll be implementing and porting the MTP module for speculative decoding next Local models really are the future of computing!
5d ago · github.com
- 3CP
2018 · github.com
- 4

- 5

Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone. - leonickson1/Swiftlet
Aug 2026 · github.com
- 6Q2
Last week was big for open source LLMs. We got: - Qwen 2.5 VL (72b and 32b) - Gemma-3 (27b) - DeepSeek-v3-0324 And a couple weeks ago we got the new mistral-ocr model. We updated our OCR benchmark to include the new models. We evaluated 1,000 documents for JSON extraction accuracy. Major takeaways: - Qwen 2.5 VL (72b and 32b) are by far the most impressive. Both landed right around 75% accuracy (equivalent to GPT-4o’s performance). Qwen 72b was only 0.4% above 32b. Within the margin of error. - Both Qwen models passed mistral-ocr (72.2%), which is specifically trained for OCR. - Gemma-3…
2025 · github.com
- 7

- 8

- 9

Private, local transcription and system-wide dictation for Apple Silicon. - VladUZH/qwen-scribe
Jul 2026 · github.com
- 10PA
2017 · github.com
- 11PA
- 12QC
2013 · zacharyvoase.com
- 13IB
2021 · kdab.com
- 14

- 15

- 16CP
2020 · github.com
- 17SL
2015 · sqlbolt.com
- 18QT
2015 · quickqwerty.com
- 19QA
2014 · github.com
- 20RC
2020 · github.com
- 21AT
2013 · dl.dropboxusercontent.com
- 22

Metal-first MoE inference for Apple Silicon with bounded SSD expert streaming and local OpenAI Chat, Responses, and Anthropic Messages endpoints. - hebrus-labs/hebrus
Jul 2026 · github.com
- 23NA
2020 · github.com
- 24IM
2013 · eggerapps.at
Ranked by how close each launch is in meaning, then by votes. Refine with a description →