nowfound

Alternatives

Products that do what Monte does

The judgment layer your AI agent is missing.

  1. 1

    Evaluate AI workflows and reach 99% AI quality.

    Oct 2025

  2. 2

    Get real-world tasks done with autonomous AI agents

    Jun 2026 · arena.ai

  3. 3

    Production-tested architecture for autonomous Claude agents

    Apr 2026 · dvdshn.com

  4. 4TB

    After training calculator agent via RL, I really wanted to go bigger! So I built RL infrastructure for training long-horizon terminal/coding agents that scales from 2x A100s to 32x H100s (~$1M worth of compute!) Without any training, my 32B agent hit #19 on Terminal-Bench leaderboard, beating Stanford's Terminus-Qwen3-235B-A22! With training... well, too expensive, but I bet the results would be good! *What I did*: - Created a Claude Code-inspired agent (system msg + tools) - Built Docker-isolated GRPO training where each rollout gets its own container - Developed a multi-agent…

    2025 · github.com

  5. 5

    Claro runs the AI agents that operate on your data

    Apr 2026 · getclaro.ai

  6. 6
    Monte1

    Portable persona for AI agents

    May 2026 · monteengine.com

  7. 7

    Your question, deliberated by 5 frontier AI models

    Apr 2026 · pilot5.ai

  8. 8

    Let your AI agent ask 1M real people before it writes

    Jul 2026 · neuroflash.com

  9. 9OS

    We implemented Stanford's Agentic Context Engineering paper which shows agents can improve their performance just by evolving their own context. How it works: Agents execute tasks, reflect on what worked/failed, and curate a "playbook" of strategies. All from execution feedback - no training data needed. Happy to answer questions about the implementation or the research!

    Oct 2025 · github.com

  10. 10AB
  11. 11

    The Only AI Tool That Doesn't Trust AI

    Mar 2026 · triall.ai

  12. 12PR

    I built this because I couldn't find honest numbers on how well VLA models [1] actually work on commercial tasks. I come from search ranking at Google where you measure everything, and in robotics nobody seemed to know. PhAIL runs four models (OpenPI/pi0.5, GR00T, ACT, SmolVLA) on bin-to-bin order picking – one of the most common warehouse operations. Same robot (Franka FR3), same objects, hundreds of blind runs. The operator doesn't know which model is running. Best model: 64 UPH. Human teleoperating the same robot: 330. Human by hand: 1,300+. Everything is public – every run with…

    Mar 2026 · phail.ai

  13. 13

    Build autonomous AI agents with Next.js & Vercel

    Oct 2025

  14. 14BA

    I'm one of the creators of The Edge Agent (TEA). We built this because we needed a way to deploy agents that was verifiable and robust enough for production/edge cases, moving away from loose scripts. The architecture aims to solve critical gaps in deterministic orchestration identified by *Prof. Claudionor Coelho Jr. (Stanford alum, ML/DL Faculty at Santa Clara Univ., and Senior Fellow for AI at Majestic Labs)* during our work on the Kiroku project. *Key Technical Features:* * *Neurosymbolic Native:* We integrated Prolog to logically validate LLM outputs. This combines neural…

    Jan 2026 · fabceolin.github.io

  15. 15MD

    We’re excited to share ML-Dev-Bench, a new open-source benchmark that tests AI agents on real-world ML development tasks. Unlike typical coding challenges or Kaggle-style competitions, our benchmark simulates end-to-end ML workflows including: - Dataset handling and preprocessing - Debugging model and code failures - Implementing new model architectures - Fine-tuning and improving existing models With 30 diverse tasks, ML-Dev-Bench evaluates agents across critical stages of ML development. To complement this, we built Calipers, a framework that provides systematic performance evaluation and…

    2025 · github.com

  16. 16
    ARA8

    Give AI systems a memory of every decision they make

    Jul 2026 · aralabs.ai

  17. 17

    Love OpenClaw? Now ship it to production. Built in Rust.

    Feb 2026 · github.com

  18. 18RS

    What relai-sdk is an open-source toolkit for making AI agents reliable via a complete learning loop: simulate → evaluate → optimize. Why Agent runs are stochastic; tool-calls fail; hard to reproduce, measure, and fix at scale. It’s also hard to align behavior with goals across output quality/format, cost, and latency. We need a loop that integrates user feedback and LLM evaluators directly into the agent code (prompts, configs, models, graphs) without overfitting. How - Simulation: LLM personas, mocked MCP servers/tools, synthetic data; can condition on real traces - Evaluation:…

    Oct 2025 · github.com

  19. 19
    Brain8

    A small, blazingly fast and extensible agent runtime

    4d ago · github.com

  20. 20

    Open-source AI agent runtime — build Agents in plain English

    Jul 2026 · syntheticbrew.ai

  21. 21

    AI Flight Simulator for Supply Chain Disruptions

    Jan 2026 · resiliencexai.com

  22. 22

    Your AI doesn't know your personality. Fix that.

    Mar 2026 · syntheticyou.com

  23. 23WB

    Hi everyone, We have been developing a platform to enable professionals to build AI assistants to help them through their work. After a few months, we realized people are trying to sell basic functionalities that can be built from scratch in a couple of hours. Due to this, individuals who are not familiar with the current SOTA are misinformed about the potential of generative models. So, we decided to open up some of our most popular templates as standalone tools for free to empower individuals and set a solid standard for what people should expect. We believe the barrier to accessing…

    2024 · join.modularmind.app

  24. 24CA

    Hi HN — I'm the creator of FastMCP and wanted to share a new project we've open-sourced called Colin. I obviously love MCP, but I also use skills extremely heavily in my day-to-day work. Being exposed to both has made me very aware of a tension: - Anything with dynamic information, I ship over MCP. This takes work to set up and requires conversational boilerplate to refresh in every conversation. - Anything behavioral, I put in skills. They're lightweight, used automatically, and feel great. But I would never put dynamic information in a skill because keeping it up to date is a pain. And yet…

    Jan 2026 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →