nowfound

Alternatives

Products that do what Eval Sakhi does

Turn your AI idea into a clear eval plan — before you ship

  1. 1
    Atla493

    Automatically detect errors in your AI agents

    Sep 2025

  2. 2
    Mukh.1206

    Automate work with AI agents

    2025

  3. 3
    EarlyAI359

    AI Agent for test code generation

    2024

  4. 4
    AgentX523

    Evaluate AI agent, pinpoint issues, and fix with one click.

    Jun 2026 · agentx.so

  5. 5AS
  6. 6
    Prefactor586

    Evaluate your AI Agents in real-time

    Jul 2026 · prefactor.tech

  7. 7
    Scorecard391

    Evaluate, Optimize, and Ship AI Agents

    Oct 2025

  8. 8
    Tessl260

    Optimize agents skills, ship 3× better code.

    Feb 2026 · tessl.io

  9. 9
    oqoqo340

    Build evals and custom benchmarks for real-world tasks

    28d ago · oqoqo.ai

  10. 10

    An open benchmark for AI agents that test APIs

    May 2026 · resources.kusho.ai

  11. 11
    Selene 1196

    Evaluate your AI app with the most accurate LLM Judge

    2025

  12. 12
    Illusion136

    Create your own AI tools, no code skills required

    2023

  13. 13

    An instantly deployable, fully customizable AI SaaS business

    2023

  14. 14AE

    I've been working on a site [1] to give people control of their LLM workflows through AI evals - automated checks that, once defined, let you move fast without regressions and cut through hype with proof. That one-liner is aimed at software engineers, but I've spent my career helping cross-functional teams collaborate, and that's really what this is about. AI agents make powerful workflows very plausible, but only if teams can grow them incrementally without losing control - no vendor lock-in, no discipline silos, no blind trust in outputs. The site tries to meet different audiences where…

    Feb 2026 · ai-evals.io

  15. 15
    Contral154

    The agentic IDE which teaches while you build.

    Mar 2026 · contral.ai

  16. 16

    Observe, evaluate and debug AI agents

    2025

  17. 17

    AI that builds you a deterministic evaluation in minutes

    2025

  18. 18
    Eval-X5

    See how engineers think with AI, not just what they build

    Jun 2026 · eval-x.com

  19. 19WE

    Hey HN! We’ve been building an MCP server to help AI-assisted web app developers by using browser agents to test whether changes made by an AI inside an editor actually work. We've been testing it on scenarios like verifying new flows in a UI, or checking that sending a chat request triggers a response. The idea is to let your coding agent both code and evaluate if what it did was correct. Here’s a short demo with Cursor: https://www.youtube.com/watch?v=_AoQK-bwR0w When building apps, we found the hardest part of AI-assisted coding isn’t the coding—it’s tedious point-and-click…

    2025 · github.com

  20. 20CL

    Hi HN! Run it: OPENROUTER_API_KEY="sk" npx bff-eval --demo We built a tool to help people take LLM outputs and easily grade them / eval them to know how good an assistant response is. We've built a number of LLM apps, and while we could ship decent tech demos, we were disappointed with how they'd perform over time. We worked with a few companies who had the same problem, and found out scientifically building prompts and evals is far from a solved problem... writing these things feels more like directing a play than coding. Inspired by Anthropic's constitutional ai concepts, and amazing…

    2025 · github.com

  21. 21CE

    Hi HN - we are the creators of “continuous-eval”, an open-source tool to test and evaluate generative AI apps. "Continuous-eval" came from our efforts to measure, validate and improve the reliability of a finance AI copilot we were developing for banks. End-to-end evaluation was not enough for us. We wanted to have granular evaluations that help pinpoint the bottlenecks and identify what / how to improve. We’ve since developed more metrics and made the framework more flexible so it can evaluate components like agent tool use, code change, retrieval steps, etc. Let us know what you think…

    2024 · github.com

  22. 22

    Hi HN! We're Giacomo and Roberto, authors of Ratel (https://github.com/ratel-ai/ratel) We used to help SaaS companies build agents on top of their products. Whenever we wanted to expand the agents’ complexity/scope, by adding more and more tools and instructions, we always run in the same issue: context bloat, with frequent hallucinations and sky high token bills. So we started constantly engineering the agents, dynamically loading tools, splitting them into subagents, inventing our own way to support skills And that's exactly when we started building Ratel: a…

    Jul 2026 · github.com

  23. 23CA
  24. 24

    Your #1 new customer is an AI agent. Are they getting an A?

    May 2026 · saastr.ai

Ranked by how close each launch is in meaning, then by votes. Refine with a description →