nowfound

Alternatives

Products that do what MiniCPM-V 4.6 does

Ultra-efficient 1.3B vision-language model for mobile

  1. 1

    GPT-4o level vision model on the phone

    2025

  2. 2

    The on-device model for your personal data

    Sep 2025

  3. 3

    Ultra-efficient on-device AI, now even faster

    2025

  4. 4

    A new SOTA for compact open models on the edge

    May 2026

  5. 5
    SmolVLM2206

    Smallest Video LM Ever from HuggingFace

    2025

  6. 6
    GLM-4.6V239

    Open-source multimodal model with native tool use

    Dec 2025

  7. 7
    NVLM 1.0200

    Open frontier-class multimodal LLMs

    2024

  8. 8
    VoxCPM2110

    Open-source 48kHz TTS with voice design and cloning

    Apr 2026

  9. 9
    InternVL3135

    Open MLLMs excelling in vision, reasoning & long context

    2025

  10. 10

    The next generation of the Phi family from Microsoft

    2025

  11. 11

    Vision-to-code foundation model for real GUI automation

    Apr 2026

  12. 12
    SmolVLA139

    Powerful robotics VLA that runs on consumer hardware

    2025

  13. 13

    The first open model to beat Sonnet made for productivity

    Feb 2026

  14. 14

    Microsoft’s New Small Language Model For Complex Reasoning

    2024

  15. 15
    Molmo 298

    SOTA video understanding, pointing, and tracking VLM

    Dec 2025

  16. 16
    Ferret193

    Refer and ground anything anywhere at any granularity

    2024

  17. 17
    GLM-5154

    Open-weights model for long-horizon agentic engineering

    Feb 2026

  18. 18

    256M VLM for end-to-end document AI

    2025

  19. 19
    Dream 7B191

    Powerful Open Diffusion LLM, Beyond Autoregressive

    2025

  20. 20

    Avoid OpenAI downtimes - one API for 30+ LLMs

    2023

  21. 21

    Open-weight 15B multimodal model for thinking and GUI agents

    Mar 2026

  22. 22

    Advanced Visual Reasoning & Agentic Tool Use

    2025

  23. 23OU

    The traditional pipeline for unstructured data extraction typically follows these steps: 1. Image → OCR Model (e.g., Google Vision) → Layout Model (e.g. Surya) → LLM → Final Answer However, this can be streamlined using a Vision-Language Model (VLM): 2. Image → VLM → Final Answer Recently VLMs have improved a lot for OCR and document understanding tasks, specifically the Qwen-2.5-VL series. We can run the Qwen-2.5-VL-7B-AWQ model locally with just 16GB VRAM, and perform end-to-end information extraction (fields and table extraction) without any external models. Hallucination with VLMs One…

    2025 · github.com

  24. 24MA

    I've been working on training this small vision language model for the last month - excited to release the first prototype today! It is based on SigLIP (image encoder), Phi-1.5 (text model) and trained using the LLaVa-1.5 training dataset. It runs reasonably fast on CPU with ~8GB of RAM in full 32-bit precision. There's plenty of room to speed it up and reduce memory consumption by quantizing the model. I posted a video of it running on my M2 Macbook Air (on CPU not MPS, so performance should be comparable on other hardware) on Twitter to demonstrate inference speed:…

    2023 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →