nowfound

Alternatives

Products that do what dutchman labs - Eval Studio does

Generate eval datasets and test your agent in minutes

  1. 1
    AgentX523

    Evaluate AI agent, pinpoint issues, and fix with one click.

    Jun 2026 · agentx.so

  2. 2AS
  3. 3WE

    Hey HN! We’ve been building an MCP server to help AI-assisted web app developers by using browser agents to test whether changes made by an AI inside an editor actually work. We've been testing it on scenarios like verifying new flows in a UI, or checking that sending a chat request triggers a response. The idea is to let your coding agent both code and evaluate if what it did was correct. Here’s a short demo with Cursor: https://www.youtube.com/watch?v=_AoQK-bwR0w When building apps, we found the hardest part of AI-assisted coding isn’t the coding—it’s tedious point-and-click…

    2025 · github.com

  4. 4
    EarlyAI359

    AI Agent for test code generation

    2024

  5. 5
    oqoqo340

    Build evals and custom benchmarks for real-world tasks

    27d ago · oqoqo.ai

  6. 6

    Test your AI apps with thousands of digital humans

    2025

  7. 7

    AI that builds you a deterministic evaluation in minutes

    2025

  8. 8

    Ship faster and test smarter with simplified AI-powered QA

    2024

  9. 9

    The visual feedback tool for AI agents

    Mar 2026 · agentation.com

  10. 10
    Atla493

    Automatically detect errors in your AI agents

    Sep 2025

  11. 11
    Agenta362

    Open-source prompt management & evals for AI teams

    Nov 2025

  12. 12PG

    I use AI agents to build UI features daily. The thing that kept annoying me: the agent writes code but never sees what it actually looks like in the browser. It can’t tell if the layout is broken or if the console is throwing errors. So I built a CLI that lets the agent open a browser, interact with the page, record what happens, and collect any errors. Then it bundles everything — video, screenshots, logs — into a self-contained HTML file I can review in seconds. proofshot start --run "npm run dev" --port 3000 # agent navigates, clicks, takes screenshots proofshot stop It works with…

    Mar 2026 · github.com

  13. 13

    The CLI your coding agent uses to ship agents

    Jul 2026 · github.com

  14. 14

    Build your AI agent in minutes to delegate the daily grind

    Dec 2025

  15. 15

    Generate no-code automation agents from screen recordings

    2024

  16. 16

    Deterministic offline release evidence for AI agents

    Jul 2026 · iisacc-justmoong.github.io

  17. 17CL

    Hi HN! Run it: OPENROUTER_API_KEY="sk" npx bff-eval --demo We built a tool to help people take LLM outputs and easily grade them / eval them to know how good an assistant response is. We've built a number of LLM apps, and while we could ship decent tech demos, we were disappointed with how they'd perform over time. We worked with a few companies who had the same problem, and found out scientifically building prompts and evals is far from a solved problem... writing these things feels more like directing a play than coding. Inspired by Anthropic's constitutional ai concepts, and amazing…

    2025 · github.com

  18. 18CE

    Hi HN - we are the creators of “continuous-eval”, an open-source tool to test and evaluate generative AI apps. "Continuous-eval" came from our efforts to measure, validate and improve the reliability of a finance AI copilot we were developing for banks. End-to-end evaluation was not enough for us. We wanted to have granular evaluations that help pinpoint the bottlenecks and identify what / how to improve. We’ve since developed more metrics and made the framework more flexible so it can evaluate components like agent tool use, code change, retrieval steps, etc. Let us know what you think…

    2024 · github.com

  19. 19

    Lets your AI agents debug production without redeploying

    7d ago · hyperprobe.co

  20. 20

    An open benchmark for AI agents that test APIs

    May 2026 · resources.kusho.ai

  21. 21AE

    I've been working on a site [1] to give people control of their LLM workflows through AI evals - automated checks that, once defined, let you move fast without regressions and cut through hype with proof. That one-liner is aimed at software engineers, but I've spent my career helping cross-functional teams collaborate, and that's really what this is about. AI agents make powerful workflows very plausible, but only if teams can grow them incrementally without losing control - no vendor lock-in, no discipline silos, no blind trust in outputs. The site tries to meet different audiences where…

    Feb 2026 · ai-evals.io

  22. 22AE

    I’ve spent the past 10 years working on AI in finance, with much of that time focused on building evaluation systems for production environments. As agents become more widely adopted, more software engineering and product people have start building them. But I’ve noticed that many teams are not yet fluent in systematic evaluation, or in the processes needed to keep agent quality high over time. For large organizations, that gap is rarely the bottleneck due to dedicated teams. But after speaking with a number of startups, it became clear that building strong, up-to-date evals is much harder…

    May 2026 · github.com

  23. 23

    Rippletide CLI is an evaluation tool for AI agents

    Jan 2026

  24. 24

    Build & run AI agents on free premium LLMs

    2025

Ranked by how close each launch is in meaning, then by votes. Refine with a description →