nowfound

Alternatives

Products that do what Qwen 3.5 running on a $300 Android phone – on-device, open source does

Qwen 3.5 Small dropped two days ago. I had it running on a mid-tier Android phone within hours. It's great seeing the on-device AI community light up around this release. Off Grid brings it to Android: phones with 6GB RAM in the $200-300 range, ~8 tok/sec on the 2B model. Fully offline. Text generation, vision AI, image gen, voice transcription, tool calling, document analysis — all on-device, nothing uploaded, ever. Works in airplane mode. 780+ GitHub stars. ~2,000 downloads across Android and iOS. Early days. GitHub:…

  1. 1RA
  2. 2

    0.8B-9B native multimodal w/ more intelligence, less compute

    Mar 2026

  3. 3
    Qwen 2.5290

    Alibaba's latest AI model series

    2025

  4. 4
    Qwen3.5307

    The 397B native multimodal agent with 17B active params

    Feb 2026

  5. 5

    Qwen's now in mobile chat

    2025

  6. 6

    SOTA open-source T2I model with even greater realism

    Jan 2026

  7. 7

    Qwen’s most capable model for coding and cowork

    Aug 2026 · qwen.ai

  8. 8Q2

    Last week was big for open source LLMs. We got: - Qwen 2.5 VL (72b and 32b) - Gemma-3 (27b) - DeepSeek-v3-0324 And a couple weeks ago we got the new mistral-ocr model. We updated our OCR benchmark to include the new models. We evaluated 1,000 documents for JSON extraction accuracy. Major takeaways: - Qwen 2.5 VL (72b and 32b) are by far the most impressive. Both landed right around 75% accuracy (equivalent to GPT-4o’s performance). Qwen 72b was only 0.4% above 32b. Within the margin of error. - Both Qwen models passed mistral-ocr (72.2%), which is specifically trained for OCR. - Gemma-3…

    2025 · github.com

  9. 9OG

    Your phone has a GPU more powerful than most 2018 laptops. Right now it sits idle while you pay monthly subscriptions to run AI on someone else's server, sending your conversations, your photos, your voice to companies whose privacy policy you've never read. Off Grid is an open-source app that puts that hardware to work. Text generation, image generation, vision AI, voice transcription — all running on your phone, all offline, nothing ever uploaded. That means you can use AI on a flight with no wifi. In a country with internet censorship. In a hospital where cloud services are a compliance…

    Feb 2026 · github.com

  10. 10

    The open sparse MoE model for agentic coding

    Apr 2026 · qwen.ai

  11. 11

    Run Qwen's latest models locally on your iPhone

    Mar 2026

  12. 12RQ
  13. 13

    A powerful open model for agentic coding tasks

    2025

  14. 14

    The end-to-end model powering multimodal chat

    2025

  15. 15

    Multimodal AI optimized for real-world coding agents

    Apr 2026 · qwen.ai

  16. 16

    The sweet-spot open dense model for coding agents

    Apr 2026 · qwen.ai

  17. 17

    A native omni model for voice, video, and tools

    Mar 2026 · qwen.ai

  18. 18
    QwQ-32B197

    Matching R1 reasoning yet 20x smaller

    2025

  19. 19

    Stunning images and perfect text

    2025

  20. 20

    Qwen's most advanced reasoning model yet

    2025

  21. 218F

    Hi HN! I'm just sharing a project I've been working on during the LLM Efficiency Challenge - you can now finetune Llama with QLoRA 5x faster than Huggingface's original implementation on your own local GPU. Some highlights: 1. Manual autograd engine - hand derived backprop steps. 2. QLoRA / LoRA 80% faster, 50% less memory. 3. All kernels written in OpenAI's Triton language. 4. 0% loss in accuracy - no approximation methods - all exact. 5. No change of hardware necessary. Supports NVIDIA GPUs since 2018+. CUDA 7.5+. 6. Flash Attention support via Xformers. 7. Supports 4bit and 16bit…

    2023 · github.com

  22. 22

    Found that you can actually run a 35B Qwen model on a Pi with very impressive intelligence and stability. Built connectors for car ODB to read all about car internals, and manufacturer's cloud service for stuff like changing AC or opening/ locking doors. Gave it info such as the full car manual. And then hooked it up with my other agents in our discussion room! So now it can answer car questions such as "when should I add oil and what kind of oil?" and help you fully offline, and when online talk with the agent family that includes all the most powerful models so they know how the car…

    12d ago · github.com

  23. 23
    Qwen3149

    Think Deeper or Act Faster

    2025

  24. 24

    Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API. - carloslfu/slotstream

    5d ago · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →