nowfound

Alternatives

Products that do what Molmo 2 does

SOTA video understanding, pointing, and tracking VLM

  1. 1

    Open robotics model that reasons in 3D before acting

    May 2026

  2. 2
    SmolVLM2206

    Smallest Video LM Ever from HuggingFace

    2025

  3. 3

    Vision-to-code foundation model for real GUI automation

    Apr 2026

  4. 4
    InternVL3135

    Open MLLMs excelling in vision, reasoning & long context

    2025

  5. 5
    V-JEPA 2198

    Meta's world model for physical world understanding

    2025

  6. 6
    GLM-4.6V239

    Open-source multimodal model with native tool use

    Dec 2025

  7. 7

    Full-Stack Platform for Training Small Language Models

    Jul 2026

  8. 8
    NVLM 1.0200

    Open frontier-class multimodal LLMs

    2024

  9. 9
    GLM-5154

    Open-weights model for long-horizon agentic engineering

    Feb 2026

  10. 10
    Ferret193

    Refer and ground anything anywhere at any granularity

    2024

  11. 11

    Ultra-efficient 1.3B vision-language model for mobile

    May 2026

  12. 12

    Multilingual, Multimodal AI from Cohere

    2025

  13. 13

    GPT-4o level vision model on the phone

    2025

  14. 14

    Advanced Visual Reasoning & Agentic Tool Use

    2025

  15. 15
    Helix140

    Bring humanoid robots to life with language and vision

    2025

  16. 16

    Bilingual ASR for dialects, code-switching, and songs

    Apr 2026

  17. 17

    AI platform for deep video understanding

    2025

  18. 18
    Wan 2.2208

    The first open MoE model for AI video generation

    2025

  19. 19
    Kimi K2.5205

    Native multimodal model with self-directed agent swarms

    Jan 2026

  20. 20OU

    The traditional pipeline for unstructured data extraction typically follows these steps: 1. Image → OCR Model (e.g., Google Vision) → Layout Model (e.g. Surya) → LLM → Final Answer However, this can be streamlined using a Vision-Language Model (VLM): 2. Image → VLM → Final Answer Recently VLMs have improved a lot for OCR and document understanding tasks, specifically the Qwen-2.5-VL series. We can run the Qwen-2.5-VL-7B-AWQ model locally with just 16GB VRAM, and perform end-to-end information extraction (fields and table extraction) without any external models. Hallucination with VLMs One…

    2025 · github.com

  21. 21
    Veo209

    Google's most powerful generative video model

    2024

  22. 22

    Fast LLMs for low-latency and high-performance workflows

    Jun 2026

  23. 23

    Google's first natively multimodal embedding model

    Mar 2026

  24. 24

    Open web agents from data to deployment

    Apr 2026

Ranked by how close each launch is in meaning, then by votes. Refine with a description →