nowfound

Alternatives

Products that do what LFM2-VL does

On-device vision, now 2x faster

  1. 1

    MoE vision-language, now easier to access

    2025

  2. 2
    LFM2.5134

    The next generation of on-device AI

    Jan 2026 · liquid.ai

  3. 3
    LFM218

    New generation of hybrid models for on-device edge AI

    2025

  4. 4

    Real-time audio conversations on-device

    Oct 2025

  5. 5TV
  6. 6

    A few days ago I found myself trying out GLM 5.2 and was really positively impressed. The capabilities and security I was getting from this LLM are similar to those I've gotten from models like Claude or GPT, and this really surprised me. But then I thought, "I wonder how it would work on a normal computer like mine," and above all, "I wonder if it would work without going into OOM on a computer like mine." So I started working with the help of agents to test this possibility. I started converting the model to int4, understanding MTP usage, and if possible implementing DSA for long context.…

    Jul 2026 · github.com

  7. 7
    SmolVLM2206

    Smallest Video LM Ever from HuggingFace

    2025

  8. 8WM

    Try it out! https://glhf.chat/ Hey HN! We’ve been working for the past few months on a website to let you easily run (almost) any open-source LLM on autoscaling GPU clusters. It’s free for now while we figure out how to price it, but we expect to be cheaper than most GPU offerings since we can run the models multi-tenant. Unlike Together AI, Fireworks, etc, we’ll run any model that the open-source vLLM project supports: we don’t have a hardcoded list. If you want a specific model or finetune, you don’t have to ask us for it: you can just paste the Hugging Face link in and…

    2024 · glhf.chat

  9. 9

    Ultra-fast 309B MoE model for coding & agents

    Dec 2025 · mimo.xiaomi.com

  10. 10SU

    Here's a project I've been working on for the last few months. It's a new (I think) algorithm, that allows to adjust smoothly - and in real time - how many calculations you'd like to do during inference of an LLM model. It seems that it's possible to do just 20-25% of weight multiplications instead of all of them, and still get good inference results. I implemented it to run on M1/M2/M3 GPU. The mmul approximation itself can be pushed to run 2x fast before the quality of output collapses. The inference speed is just a bit faster than Llama.cpp's, because the rest of implementation…

    2024 · asciinema.org

  11. 11
    Qwen3.5307

    The 397B native multimodal agent with 17B active params

    Feb 2026 · qwen.ai

  12. 12
    InternVL3135

    Open MLLMs excelling in vision, reasoning & long context

    2025

  13. 13
    Molmo 298

    SOTA video understanding, pointing, and tracking VLM

    Dec 2025 · allenai.org

  14. 14

    Massive local model speedup on Apple Silicon with MLX

    Apr 2026 · ollama.com

  15. 15
    SmolVLA139

    Powerful robotics VLA that runs on consumer hardware

    2025

  16. 16

    The open-source era of 1M context intelligence

    Apr 2026 · huggingface.co

  17. 17

    Vision-to-code foundation model for real GUI automation

    Apr 2026 · docs.z.ai

  18. 18

    Open-source stack for industrial-grade LLM applications

    2025

  19. 19
    NVLM 1.0200

    Open frontier-class multimodal LLMs

    2024

  20. 20

    The Sweet Spot for Open-Source Multimodal AI

    2025

  21. 21

    Ultra-efficient 1.3B vision-language model for mobile

    May 2026 · github.com

  22. 22
    GLM-5154

    Open-weights model for long-horizon agentic engineering

    Feb 2026 · z.ai

  23. 23
    MiMo116

    Xiaomi's Open Source Model, Born for Reasoning

    2025

  24. 24
    GLM-4.5298

    Unifying agentic capabilities in one open model

    2025

Ranked by how close each launch is in meaning, then by votes. Refine with a description →