nowfound

Alternatives

Products that do what GLM-4.6V does

Open-source multimodal model with native tool use

  1. 1

    Vision-to-code foundation model for real GUI automation

    Apr 2026

  2. 2
    GLM-5154

    Open-weights model for long-horizon agentic engineering

    Feb 2026

  3. 3
    NVLM 1.0200

    Open frontier-class multimodal LLMs

    2024

  4. 4
    InternVL3135

    Open MLLMs excelling in vision, reasoning & long context

    2025

  5. 5
    Ferret193

    Refer and ground anything anywhere at any granularity

    2024

  6. 6
    GLM-4.720

    Advanced coding & reasoning with multi-turn thinking

    Dec 2025

  7. 7

    Ultra-efficient 1.3B vision-language model for mobile

    May 2026

  8. 8

    Open-weight 15B multimodal model for thinking and GUI agents

    Mar 2026

  9. 9

    GPT-4o level vision model on the phone

    2025

  10. 10
    Kimi K2.5205

    Native multimodal model with self-directed agent swarms

    Jan 2026

  11. 11

    OpenAI most advanced image generator yet

    2025

  12. 12
    Openlit152

    One click observability & evals for LLMs & GPUs

    2024

  13. 13

    Fast multimodal-native inference at scale

    Dec 2025

  14. 14

    The end-to-end model powering multimodal chat

    2025

  15. 15
    Molmo 298

    SOTA video understanding, pointing, and tracking VLM

    Dec 2025

  16. 16
    Wan 2.6151

    The next era of multimodal AI for creators is here

    Dec 2025

  17. 17
    Grok-1239

    Open source release of xAI's LLM

    2024

  18. 18
    Dream 7B191

    Powerful Open Diffusion LLM, Beyond Autoregressive

    2025

  19. 19
    SmolVLM2206

    Smallest Video LM Ever from HuggingFace

    2025

  20. 20

    7B open model mixing transformers and linear RNNs

    Mar 2026

  21. 21

    The next generation of the Phi family from Microsoft

    2025

  22. 22

    Microsoft’s New Small Language Model For Complex Reasoning

    2024

  23. 23

    Multilingual, Multimodal AI from Cohere

    2025

  24. 24
    Fuyu-8B114

    A multimodal architecture for AI agents

    2023

Ranked by how close each launch is in meaning, then by votes. Refine with a description →