nowfound

Alternatives

Products that do what FastVLM does

Lightweight open-source vision models for Apple devices

  1. 1
    Ferret193

    Refer and ground anything anywhere at any granularity

    2024

  2. 2

    MoE vision-language, now easier to access

    2025

  3. 3
    InternVL3135

    Open MLLMs excelling in vision, reasoning & long context

    2025

  4. 4

    Vision-to-code foundation model for real GUI automation

    Apr 2026 · docs.z.ai

  5. 5
    GLM-4.6V239

    Open-source multimodal model with native tool use

    Dec 2025 · z.ai

  6. 6
    SmolVLM2206

    Smallest Video LM Ever from HuggingFace

    2025

  7. 7

    Massive local model speedup on Apple Silicon with MLX

    Apr 2026 · ollama.com

  8. 8
    Molmo 298

    SOTA video understanding, pointing, and tracking VLM

    Dec 2025 · allenai.org

  9. 9

    AI image tool that lets you make edits by describing them

    2024

  10. 10

    Ultra-efficient 1.3B vision-language model for mobile

    May 2026 · github.com

  11. 11
    LFM2-VL19

    On-device vision, now 2x faster

    2025

  12. 12
    NVLM 1.0200

    Open frontier-class multimodal LLMs

    2024

  13. 13

    Take raw photos with proof they're real, not AI

    May 2026 · vwfndr.camera

  14. 14
    GLM-5154

    Open-weights model for long-horizon agentic engineering

    Feb 2026 · z.ai

  15. 15
    SmolVLA139

    Powerful robotics VLA that runs on consumer hardware

    2025

  16. 16
    FastCut362

    Create engaging captions for your short-form video with AI

    2023

  17. 17

    GPT-4o level vision model on the phone

    2025

  18. 18BV

    Vision models have been gaining popularity as a replacement for traditional OCR. Especially with Gemini 2.0 becoming cost competitive with the cloud platforms. We've been continuously evaluating different models since we released the Zerox package last year (https://github.com/getomni-ai/zerox). And we wanted to put some numbers behind it. So we’re open sourcing our internal OCR benchmark + evaluation datasets. Full writeup + data explorer here: https://getomni.ai/ocr-benchmark Github: https://github.com/getomni-ai/benchmark Huggingface:…

    2025 · getomni.ai

  19. 19TV
  20. 20
    Apple MLX138

    An array framework for machine learning on Apple silicon

    2023

  21. 21

    Immersive 3D image generation for Apple Vision Pro

    2024

  22. 22WM

    We wrote our inference engine on Rust, it is faster than llama cpp in all of the use cases. Your feedback is very welcomed. Written from scratch with idea that you can add support of any kernel and platform.

    2025 · github.com

  23. 23LA

    A simple mobile web app inspired by Fuzzy-Search/realtime-bakllava that uses llama.cpp server backend with multimodal mode to describe and narrate what the phone camera sees. I built this thing in a few hours using a single ChatGPT thread to generate most things for me and iterate on this project. Here's the workflow: https://chat.openai.com/share/ea84ec69-5617-45e8-8772-ac2dcf...

    2023 · github.com

  24. 24

    Vision encoder setting new standards in image & video tasks

    2025

Ranked by how close each launch is in meaning, then by votes. Refine with a description →