nowfound

Alternatives

Products that do what Chat with Orion – a visual agent that sees, reasons and acts does

Hey HN! We’re excited to share Orion [1] — our new visual agent that sees, reasons, and acts across images, videos, and documents. Frontier VLMs (GPT, Claude, Gemini) can describe what they see, but they can’t reliably act on visual inputs. Ask them to detect objects, segment images, or chain visual steps — they’ll fail in surprisingly inconsistent ways. High-res images collapse to ~1024px. And the visual AI ecosystem is fragmented across separate APIs for image understanding, OCR, image-gen, video-gen, etc. We built Orion to fix this. Orion combines VLM reasoning with reliable…

  1. 1

    Visual AI Agent for interacting with image, document, video

    Nov 2025

  2. 2

    Agentic visual reasoning with code execution

    Jan 2026 · blog.google

  3. 3

    Search by Seeing Instantly with Visual Reasoning Model

    2025

  4. 4

    Reasoning-Driven Agentic Object Detection

    2025

  5. 5
    Qwen3.5307

    The 397B native multimodal agent with 17B active params

    Feb 2026 · qwen.ai

  6. 6
    Gemini529

    Google's answer to GPT-4

    2023

  7. 7
    SIMA 2204

    Google's most capable AI agent for virtual 3D worlds

    Nov 2025

  8. 8

    MoE vision-language, now easier to access

    2025

  9. 9

    Your vision AI agent to build VisionAI applications

    2025

  10. 10

    Google's new AI model for the agentic era

    2024

  11. 11

    Google's SOTA robotics model for visual & spatial reasoning!

    Apr 2026 · deepmind.google

  12. 12MR

    Data visualizations are the bridge between user and data. But building AI agents that can generate visualizations reliably can be very tricky: - simple chart specs can be reliable, but generated charts are often of low quality due to reliance on system defaults; - complex chart specs with explicit details can produce good-looking charts, but they are verbose and agents can struggle with reliability We figured out it is a limitation on the language issue (not just AI capability thing) -- current visualization languages are a bit too low-level for AI agents, requiring them to explicitly make…

    Jul 2026 · microsoft.github.io

  13. 13IR
  14. 14

    Turn any LLM into a Computer Use Agent

    2025

  15. 15

    Curiosity Lens: Your Visual Agent

    2025

  16. 16TV

    Hey HN! I built a tool that gives LLMs the ability to understand the visual structure of a webpage even if they don't accept image input. We've found that unimodal GPT-4 + Tarsier's textual webpage representation consistently beats multimodal GPT-4V/4o + webpage screenshot by 10-20%, probably because multimodal LLMs still aren't as performant as they're hyped to be. Over the course of experimenting with pruned HTML, accessibility trees, and other perception systems for web agents, we've iterated on Tarsier's components to maximize downstream agent/codegen performance. Here's the…

    2024 · github.com

  17. 17

    Google’s best model for logical thinking and understanding

    Dec 2025

  18. 18
    Sun110

    Collaborative voice API for agents

    Jun 2026 · getsun.io

  19. 19VA

    Rather than calling tools one by one, Orion 2 generates a program and runs it end to end, meaning fewer round-trips and lower latency. You can try it out at https://chat.vlm.run When orchestration is code, every workflow is composable, inspectable, and deterministic. We put together a short Orion 2 demo video: https://www.youtube.com/watch?v=rzhXcNAYQ-0

    Jun 2026 · vlm.run

  20. 20BV

    Vision models have been gaining popularity as a replacement for traditional OCR. Especially with Gemini 2.0 becoming cost competitive with the cloud platforms. We've been continuously evaluating different models since we released the Zerox package last year (https://github.com/getomni-ai/zerox). And we wanted to put some numbers behind it. So we’re open sourcing our internal OCR benchmark + evaluation datasets. Full writeup + data explorer here: https://getomni.ai/ocr-benchmark Github: https://github.com/getomni-ai/benchmark Huggingface:…

    2025 · getomni.ai

  21. 21

    Advanced Visual Reasoning & Agentic Tool Use

    2025

  22. 22
    Orion2

    Your AI workspace for research, agents & workflows

    13d ago · download-orion-ai.web.app

  23. 23

    The open source agentic BI platform

    27d ago · orion-agent.ai

  24. 24
    InternVL3135

    Open MLLMs excelling in vision, reasoning & long context

    2025

Ranked by how close each launch is in meaning, then by votes. Refine with a description →