nowfound

Alternatives

Products that do what AI Evaluation Tools does

Test LLMs, RAG & AI Agents with Confidence

  1. 1

    AI prompt engineering business model guide

    2023

  2. 2

    Discover the best AI tools

    2024

  3. 3
    Okareo127

    Error discovery & evaluation for AI Agents

    2025

  4. 4

    Observe, evaluate and debug AI agents

    2025

  5. 5
    Selene 1196

    Evaluate your AI app with the most accurate LLM Judge

    2025

  6. 6

    Test-driven development for LLMs

    2023

  7. 7
    Stax179

    Move your LLM evals from vibes to data

    2025

  8. 8

    Get smarter with 100 AI-generated thinking tools

    2022

  9. 9

    Validate, monitor, and safeguard LLM-based apps

    2023

  10. 10

    LLMs price comparison tool developed and updated by LLM

    2024

  11. 11

    Improve your LLM apps with open-source observability tool

    2024

  12. 12

    An open benchmark for AI agents that test APIs

    May 2026

  13. 13

    Let AI score your translation work

    2025

  14. 14

    Aggregate uptime monitoring across OpenAI, Claude, and more

    Apr 2026

  15. 15

    Simplify your coding journey

    2024

  16. 16
    Colossal135

    Effortlessly integrate tool-using agents with a single fetch

    2025

  17. 17

    Rippletide CLI is an evaluation tool for AI agents

    Jan 2026

  18. 18

    Define tools once for agents use them everywhere

    Mar 2026

  19. 19
    ADK-TS108

    Build smart, tool-using agents in just one line

    2025

  20. 20

    Discover the best AI tools & websites

    2023

  21. 21AE

    I've been working on a site [1] to give people control of their LLM workflows through AI evals - automated checks that, once defined, let you move fast without regressions and cut through hype with proof. That one-liner is aimed at software engineers, but I've spent my career helping cross-functional teams collaborate, and that's really what this is about. AI agents make powerful workflows very plausible, but only if teams can grow them incrementally without losing control - no vendor lock-in, no discipline silos, no blind trust in outputs. The site tries to meet different audiences where…

    Feb 2026 · ai-evals.io

  22. 22

    Expand eval coverage & use red agents to break AI systems

    24d ago · mutant.aiankit.com

  23. 23NH

    Hey HN! When I started looking into LLMs and agents for software development and introducing them at work, I quickly realised that a person new to the topic faces a real barrage: - all the hype (AGI, engineers getting replaced by AI etc.) - conflicting opinions in virtually every discussion—for every person saying they’ve 10x-ed their productivity, there is a comment decrying LLMs as an utter failure - a lot of jargon (MoE, MCP, RAG, distillation, quantisation etc. etc.) - a profusion of models, IDEs/IDE extensions, CLI agents, other tools etc. Sorting through all of this can be quite…

    2025 · nohypeai.dev

  24. 24SE

    Hey HN! I built self-driving sim and eval at Waymo. Now I’m building Scorecard to bring that approach to agent eval: reproducible, automated scoring for AI. Scorecard lets you: - Run LLM-as-judge evals on agent workflows: test tool usage, multi-step reasoning, and task completion in CI/CD or in a playground. - Debug failures with OpenTelemetry traces: see which tool failed, why your agent looped, and where reasoning went wrong. - Collaborate on datasets, simulated agents, and evaluation metrics. Try it out → https://app.scorecard.io (free tier, no payment required!) Docs →…

    Oct 2025 · docs.scorecard.io

Ranked by how close each launch is in meaning, then by votes. Refine with a description →