nowfound

Alternatives

Products that do what EvalTrim does

Prove which AI-agent evals are worth keeping.

  1. 1
    Prefactor586

    Evaluate your AI Agents in real-time

    Jul 2026 · prefactor.tech

  2. 2
    oqoqo340

    Build evals and custom benchmarks for real-world tasks

    27d ago · oqoqo.ai

  3. 3

    An open benchmark for AI agents that test APIs

    May 2026

  4. 4

    Trace, evaluate, and improve AI agents in production

    Aug 2026 · telerik.com

  5. 5
    Retrace101

    Debug AI agents by replaying and forking runs

    Jul 2026 · retraceai.tech

  6. 6
    TryCase219

    Disposable test environments for AI coding agents

    Jul 2026 · trycase.dev

  7. 7

    Deterministic offline release evidence for AI agents

    Jul 2026 · iisacc-justmoong.github.io

  8. 8
    AEVS131

    proof-of-execution for AI agents

    Jun 2026 · aevs.fetch.ai

  9. 9

    The modern standard in AML compliance through AI agents

    2023

  10. 10AE

    I've been working on a site [1] to give people control of their LLM workflows through AI evals - automated checks that, once defined, let you move fast without regressions and cut through hype with proof. That one-liner is aimed at software engineers, but I've spent my career helping cross-functional teams collaborate, and that's really what this is about. AI agents make powerful workflows very plausible, but only if teams can grow them incrementally without losing control - no vendor lock-in, no discipline silos, no blind trust in outputs. The site tries to meet different audiences where…

    Feb 2026 · ai-evals.io

  11. 11

    o3 for Lawyers, AI powered Legal Research tool

    2025

  12. 12AE

    I’ve spent the past 10 years working on AI in finance, with much of that time focused on building evaluation systems for production environments. As agents become more widely adopted, more software engineering and product people have start building them. But I’ve noticed that many teams are not yet fluent in systematic evaluation, or in the processes needed to keep agent quality high over time. For large organizations, that gap is rarely the bottleneck due to dedicated teams. But after speaking with a number of startups, it became clear that building strong, up-to-date evals is much harder…

    May 2026 · github.com

  13. 13AB

    Hi everyone! My team and I just open-sourced a bunch of cool agent dev tools: Invariant Explorer to visually inspect and understand AI traces and a testing framework, building on pytest.

    2024 · github.com

  14. 14OS

    Hey HN! We built EvalKit, a library you embed to capture agent actions and a UI where domain experts give feedback, evaluate and improve AI agents. We experienced, in large agentic systems, prompt-engineering or auto-prompt improvement tool can get accuracy from 0 to 50% but for increasing accuracy to 100% we had to work with domain experts. Example -> In a law ai agent, lawyers are needed because law is complex and lawyers have a deeper context compared to non-lawyers. Other evaluation tools in the market focus on the experience of the developer and we are focusing on making as easy as…

    2025 · github.com

  15. 15

    Snapshot-test AI behavior in CI

    Jul 2026 · evalcore.cc

  16. 16BE

    Hey HN, We're excited to introduce Braintrust, a platform for running and tracking AI evaluations (“evals”) [1]. At my previous startup Impira and leading AI at Figma, we had this recurring problem where we never knew if changes we made to our products would improve or regress key user scenarios. We built some tooling to solve this problem and after talking to other developers learned that it was a widespread issue. Specifically, it’s challenging to establish a great dev loop that lets you systematically improve and ship high quality AI products. We worked with the teams at Zapier, Coda, and…

    2023

  17. 17

    Trajectory regression testing for AI agents

    17d ago · agentdiff.lostmartian.in

  18. 18

    Stop AI agents from installing malicious packages.

    Jul 2026 · agentinel.habitwala.in

  19. 19PL

    Library makes requests asynchronously across models, so you can spend a lot of $$ quickly if you want XD. But seriously I hope this enables folks to create and run evals (especially safety ones) a lot easier than before.

    2024 · github.com

  20. 20AT

    Hi Hacker News! We're launching Zalor, an agent testing platform. Agents often break when you tweak system prompts, swap models, or add tools. Zalor automatically generates test scenarios and evaluates your agent so you know it's reliable before deploying to production. We currently support the OpenAI Agents SDK and are onboarding other frameworks. A GitHub integration is coming so you can get feedback on every update. Looking forward to hearing feedback from people building agents.

    Mar 2026 · agents.zalor.ai

  21. 21

    Proof-carrying workflows for coding agents

    Jul 2026 · github.com

  22. 22

    On-brand AI work that has to prove it is ready to ship

    Jul 2026 · corristonconsulting.com

  23. 23

    Production failures become regression tests for AI agents

    26d ago · tracely-ai.com

  24. 24

    AI agent spend firewall

    25d ago · agentshield.fly.dev

Ranked by how close each launch is in meaning, then by votes. Refine with a description →