nowfound

Alternatives

Products that do what Eval-X does

See how engineers think with AI, not just what they build

  1. 1
    AgentX523

    Evaluate AI agent, pinpoint issues, and fix with one click.

    Jun 2026 · agentx.so

  2. 2

    AI that builds you a deterministic evaluation in minutes

    2025

  3. 3AS
  4. 4

    Enabling Claude to stop and think

    2025

  5. 5
    oqoqo340

    Build evals and custom benchmarks for real-world tasks

    27d ago · oqoqo.ai

  6. 6

    Prepare for any job interview with our AI interviewer

    2024

  7. 7

    All-in-one AI workspace: content, code, design & API power

    2025

  8. 8

    Helps you write better prompts

    2022

  9. 9
    Stax179

    Move your LLM evals from vibes to data

    2025

  10. 10

    AI-native engineer hiring via real-work simulations

    Dec 2025

  11. 11

    Turn any AI into a senior PM, engineer, or analyst

    Jun 2026 · mohitagw15856.github.io

  12. 12AE

    I've been working on a site [1] to give people control of their LLM workflows through AI evals - automated checks that, once defined, let you move fast without regressions and cut through hype with proof. That one-liner is aimed at software engineers, but I've spent my career helping cross-functional teams collaborate, and that's really what this is about. AI agents make powerful workflows very plausible, but only if teams can grow them incrementally without losing control - no vendor lock-in, no discipline silos, no blind trust in outputs. The site tries to meet different audiences where…

    Feb 2026 · ai-evals.io

  13. 13

    Gemini-powered AI agent for insights & faster ad decisions

    Jun 2026 · blog.google

  14. 14CL

    Hi HN! Run it: OPENROUTER_API_KEY="sk" npx bff-eval --demo We built a tool to help people take LLM outputs and easily grade them / eval them to know how good an assistant response is. We've built a number of LLM apps, and while we could ship decent tech demos, we were disappointed with how they'd perform over time. We worked with a few companies who had the same problem, and found out scientifically building prompts and evals is far from a solved problem... writing these things feels more like directing a play than coding. Inspired by Anthropic's constitutional ai concepts, and amazing…

    2025 · github.com

  15. 15MA

    I've been exploring the (not so=) amazing potential of AI in coding and have compiled a list of tools. From AI-powered IDEs to code generators, this resource is my contribution to the community. I'm still on the fence about including txt2sql projects, as their functionality seems too basic to me. And I'm personally maintaining this, so your feedback is wellcome.

    2025 · aicode.danvoronov.com

  16. 16AE

    I’ve spent the past 10 years working on AI in finance, with much of that time focused on building evaluation systems for production environments. As agents become more widely adopted, more software engineering and product people have start building them. But I’ve noticed that many teams are not yet fluent in systematic evaluation, or in the processes needed to keep agent quality high over time. For large organizations, that gap is rarely the bottleneck due to dedicated teams. But after speaking with a number of startups, it became clear that building strong, up-to-date evals is much harder…

    May 2026 · github.com

  17. 17

    Let AI score your translation work

    2025

  18. 18IM

    This is another one of my automate-my-life projects - I'm constantly asking the same question to different AIs since there's always the hope of getting a better answer somewhere else. Maybe ChatGPT's answer is too short, so I ask Perplexity. But I realize that's hallucinated, so I try Gemini. That answer sounds right, but I cross-reference with Claude just to make sure. This doesn't really apply to math/coding (where o1 or Gemini can probably one-shot an excellent response), but more to online search, where information is more fluid and there's no "right" search engine + text…

    2024 · ithy.com

  19. 19

    o3 for Lawyers, AI powered Legal Research tool

    2025

  20. 20BE

    Hey HN, We're excited to introduce Braintrust, a platform for running and tracking AI evaluations (“evals”) [1]. At my previous startup Impira and leading AI at Figma, we had this recurring problem where we never knew if changes we made to our products would improve or regress key user scenarios. We built some tooling to solve this problem and after talking to other developers learned that it was a widespread issue. Specifically, it’s challenging to establish a great dev loop that lets you systematically improve and ship high quality AI products. We worked with the teams at Zapier, Coda, and…

    2023

  21. 21

    Hire 20× Faster. Fill Roles in Days, Not Weeks

    Aug 2026 · justinterview.ai

  22. 22CE

    Hi HN - we are the creators of “continuous-eval”, an open-source tool to test and evaluate generative AI apps. "Continuous-eval" came from our efforts to measure, validate and improve the reliability of a finance AI copilot we were developing for banks. End-to-end evaluation was not enough for us. We wanted to have granular evaluations that help pinpoint the bottlenecks and identify what / how to improve. We’ve since developed more metrics and made the framework more flexible so it can evaluate components like agent tool use, code change, retrieval steps, etc. Let us know what you think…

    2024 · github.com

  23. 23

    Stop trusting vendor demos. Interrogate them with AI first.

    Apr 2026 · github.com

  24. 24PE

    Spelltest framework simulates conversations between AI ‘synthetic users' in an environment to test and refine LLM-based applications. It ensures your app converse with utmost accuracy and relevance. Post-chat, Spelltest assesses responses, providing qualitative and quantitative feedback on performance. Suitable for both chat and completion modes. When to use: - After modifying your prompt. - When your LLM provider updates. - As a CI step for you repo. All feedback and collaborations appreciated!

    2023 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →