nowfound

Alternatives

Products that do what Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone does

  1. 1RQ
  2. 2Q6
  3. 3

    0.8B-9B native multimodal w/ more intelligence, less compute

    Mar 2026

  4. 4

    Qwen’s most capable model for coding and cowork

    Aug 2026 · qwen.ai

  5. 5

    Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API. - carloslfu/slotstream

    5d ago · github.com

  6. 6
    Qwen 2.5290

    Alibaba's latest AI model series

    2025

  7. 7
    Qwen3.5307

    The 397B native multimodal agent with 17B active params

    Feb 2026

  8. 8OS

    Hi HN, I built a specialized inference engine for running 4-bit Gemma 4 26B-A4B-IT on any M-series Mac using about 2 GB of RAM. It is called TurboFieldfare and is written in Swift and Metal. I have always adored on-device AI. It feels like magic that you can run a powerful NN on your Mac or iPhone. So I wanted to push the limits a bit and run a model whose weights don’t fit in memory. The model’s 4-bit quantized weights occupy roughly 14 GB, which makes running it with conventional inference tools almost impossible on an 8 GB or even 16 GB Mac once the OS, applications, and KV cache are…

    Jul 2026 · github.com

  9. 9EL
  10. 10

    Run Qwen's latest models locally on your iPhone

    Mar 2026

  11. 11

    Large language model series developed by Alibaba Cloud

    2025

  12. 12

    Qwen's most advanced reasoning model yet

    2025

  13. 13

    How small can a language model be while still doing something useful? I wanted to find out, and had some spare time over the holidays. Z80-μLM is a character-level language model with 2-bit quantized weights ({-2,-1,0,+1}) that runs on a Z80 with 64KB RAM. The entire thing: inference, weights, chat UI, it all fits in a 40KB .COM file that you can run in a CP/M emulator and hopefully even real hardware! It won't write your emails, but it can be trained to play a stripped down version of 20 Questions, and is sometimes able to maintain the illusion of having simple but terse conversations…

    Dec 2025 · github.com

  14. 14

    The end-to-end model powering multimodal chat

    2025

  15. 15

    A powerful open model for agentic coding tasks

    2025

  16. 16

    The open sparse MoE model for agentic coding

    Apr 2026 · qwen.ai

  17. 17FT

    Aug 2026 · github.com

  18. 18MP
  19. 19
    QwQ-32B197

    Matching R1 reasoning yet 20x smaller

    2025

  20. 20Q2

    Last week was big for open source LLMs. We got: - Qwen 2.5 VL (72b and 32b) - Gemma-3 (27b) - DeepSeek-v3-0324 And a couple weeks ago we got the new mistral-ocr model. We updated our OCR benchmark to include the new models. We evaluated 1,000 documents for JSON extraction accuracy. Major takeaways: - Qwen 2.5 VL (72b and 32b) are by far the most impressive. Both landed right around 75% accuracy (equivalent to GPT-4o’s performance). Qwen 72b was only 0.4% above 32b. Within the margin of error. - Both Qwen models passed mistral-ocr (72.2%), which is specifically trained for OCR. - Gemma-3…

    2025 · github.com

  21. 21

    Multimodal AI optimized for real-world coding agents

    Apr 2026 · qwen.ai

  22. 22

    SOTA open-source T2I model with even greater realism

    Jan 2026

  23. 23

    A native omni model for voice, video, and tools

    Mar 2026 · qwen.ai

  24. 24GG

    A few days ago I found myself trying out GLM 5.2 and was really positively impressed. The capabilities and security I was getting from this LLM are similar to those I've gotten from models like Claude or GPT, and this really surprised me. But then I thought, "I wonder how it would work on a normal computer like mine," and above all, "I wonder if it would work without going into OOM on a computer like mine." So I started working with the help of agents to test this possibility. I started converting the model to int4, understanding MTP usage, and if possible implementing DSA for long context.…

    Jul 2026 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →