nowfound

Alternatives

Products that do what MiniCPM-o 4.5 does

Real-time, full-duplex multimodal AI on your device

  1. 1
    GPT-Live211

    Full-duplex voice for ChatGPT

    Jul 2026 · openai.com

  2. 2

    Run leading vision models locally with the new engine

    2025

  3. 3

    GPT-4o level vision model on the phone

    2025

  4. 4

    Ultra-efficient 1.3B vision-language model for mobile

    May 2026 · github.com

  5. 5

    The on-device model for your personal data

    Sep 2025

  6. 6
    GPT-41,161

    LLM that exhibits human-level performance

    2023 · openai.com

  7. 7

    A new SOTA for compact open models on the edge

    May 2026 · huggingface.co

  8. 8

    Run multimodal AI locally with an encoder-free architecture

    Jun 2026 · blog.google

  9. 9

    Build Powerful Voice Agents

    2025

  10. 10

    Tighter instruction adherence in speech agents

    Feb 2026 · developers.openai.com

  11. 11

    Ultra-efficient on-device AI, now even faster

    2025

  12. 12

    For reliable, production-ready voice agents

    2025

  13. 13
    Gemma 3n199

    Run powerful multimodal AI right on your phone

    2025

  14. 14

    The end-to-end model powering multimodal chat

    2025

  15. 15

    The next generation of the Phi family from Microsoft

    2025

  16. 16

    MiniMax H3 (aka Hailuo 3) is a multimodal AI video generator. Turn text and images into cinematic clips with native audio and omni-reference — free to start.

    Jul 2026 · minimax.io

  17. 17
    Llama 4423

    A new era of natively multimodal AI innovation

    2025

  18. 18
    Llama312

    3.1-405B: an open source model to rival GPT-4o / Claude-3.5

    2024

  19. 19

    Massive local model speedup on Apple Silicon with MLX

    Apr 2026 · ollama.com

  20. 20TN

    Kitten TTS (https:&#x2F;&#x2F;github.com&#x2F;KittenML&#x2F;KittenTTS) is an open-source series of tiny and expressive text-to-speech models for on-device applications. We had a thread last year here: https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=44807868. Today we're releasing three new models with 80M, 40M and 14M parameters. The largest model (80M) has the highest quality. The 14M variant reaches new SOTA in expressivity among similar sized models, despite being <25MB in size. This release is a major upgrade from the previous one and supports English text-to-speech applications in…

    Mar 2026 · github.com

  21. 21

    Advanced Visual Reasoning & Agentic Tool Use

    2025

  22. 22
    Mistral 3415

    A family of frontier open-source multimodal models

    Dec 2025 · mistral.ai

  23. 23

    0.8B-9B native multimodal w/ more intelligence, less compute

    Mar 2026 · huggingface.co

  24. 24
    VoxCPM2110

    Open-source 48kHz TTS with voice design and cloning

    Apr 2026 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →