Alternatives
Products that do what Qwen3.6-35B-A3B on a 16 GB M1 Pro with SSD-streamed MoE does
- 1
Qwen3.5 Small▲3590.8B-9B native multimodal w/ more intelligence, less compute
Mar 2026 · huggingface.co
- 2

- 3

- 4

- 5

- 6RQ
Sep 2025 · github.com
- 7

- 8

- 9

Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone. - leonickson1/Swiftlet
Aug 2026 · github.com
- 10

I built slotstream, a way to run Qwen3.8-Flash-Next 4-bit on a low-memory mac starting from 16GB, a 125B parameter model that would need 100GB+ memory/RAM, thanks to expert-offloading/ssd-streaming. Easy to install/update, and mac-native using MLX and Swift. It ships with auto-mode, which makes a good tradeoff between memory usage and speed. I'll be implementing and porting the MTP module for speculative decoding next Local models really are the future of computing!
5d ago · github.com
- 11

- 12

- 13

- 14

- 15

- 16

- 17Q2
Last week was big for open source LLMs. We got: - Qwen 2.5 VL (72b and 32b) - Gemma-3 (27b) - DeepSeek-v3-0324 And a couple weeks ago we got the new mistral-ocr model. We updated our OCR benchmark to include the new models. We evaluated 1,000 documents for JSON extraction accuracy. Major takeaways: - Qwen 2.5 VL (72b and 32b) are by far the most impressive. Both landed right around 75% accuracy (equivalent to GPT-4o’s performance). Qwen 72b was only 0.4% above 32b. Within the margin of error. - Both Qwen models passed mistral-ocr (72.2%), which is specifically trained for OCR. - Gemma-3…
2025 · github.com
- 18

- 19

- 20

- 21

- 22

- 23

- 24MP
Aug 2026 · deepgrove.ai
Ranked by how close each launch is in meaning, then by votes. Refine with a description →