nowfound

Alternatives

Products that do what Cipherra does

Infrastructure for Continuous Evals of AI Agents

  1. 1
    AgentX523

    Evaluate AI agent, pinpoint issues, and fix with one click.

    Jun 2026 · agentx.so

  2. 2
    oqoqo340

    Build evals and custom benchmarks for real-world tasks

    27d ago · oqoqo.ai

  3. 3
    Atla493

    Automatically detect errors in your AI agents

    Sep 2025

  4. 4

    AI agents that find, validate, and fix every vulnerability

    Jun 2026 · getastra.com

  5. 5

    Agentic testing for the AI-native team.

    Mar 2026

  6. 6

    An open benchmark for AI agents that test APIs

    May 2026 · resources.kusho.ai

  7. 7

    Parallel custom agents for complex tasks

    Mar 2026 · learn.chatgpt.com

  8. 8

    Get real-world tasks done with autonomous AI agents

    Jun 2026 · arena.ai

  9. 9AS
  10. 10TB

    After training calculator agent via RL, I really wanted to go bigger! So I built RL infrastructure for training long-horizon terminal/coding agents that scales from 2x A100s to 32x H100s (~$1M worth of compute!) Without any training, my 32B agent hit #19 on Terminal-Bench leaderboard, beating Stanford's Terminus-Qwen3-235B-A22! With training... well, too expensive, but I bet the results would be good! *What I did*: - Created a Claude Code-inspired agent (system msg + tools) - Built Docker-isolated GRPO training where each rollout gets its own container - Developed a multi-agent…

    2025 · github.com

  11. 11

    Validate every PR with AI that runs tests for you

    Apr 2026 · qa.tech

  12. 12

    The voice AI auto-testing loop to simulate, evaluate & ship

    2025

  13. 13WE

    Hey HN! We’ve been building an MCP server to help AI-assisted web app developers by using browser agents to test whether changes made by an AI inside an editor actually work. We've been testing it on scenarios like verifying new flows in a UI, or checking that sending a chat request triggers a response. The idea is to let your coding agent both code and evaluate if what it did was correct. Here’s a short demo with Cursor: https://www.youtube.com/watch?v=_AoQK-bwR0w When building apps, we found the hardest part of AI-assisted coding isn’t the coding—it’s tedious point-and-click…

    2025 · github.com

  14. 14

    Deploy AI apps instantly with a single shot

    2025

  15. 15

    Build & scale AI \ agents as microservices with IAM

    Dec 2025

  16. 16CE

    Hi HN - we are the creators of “continuous-eval”, an open-source tool to test and evaluate generative AI apps. "Continuous-eval" came from our efforts to measure, validate and improve the reliability of a finance AI copilot we were developing for banks. End-to-end evaluation was not enough for us. We wanted to have granular evaluations that help pinpoint the bottlenecks and identify what / how to improve. We’ve since developed more metrics and made the framework more flexible so it can evaluate components like agent tool use, code change, retrieval steps, etc. Let us know what you think…

    2024 · github.com

  17. 17EY

    I built an open-source AI agent for security testing to find and fix vulnerabilities in your code. I’ve noticed how bad security vulnerabilities have gotten with everyone shipping AI code slop, so I wanted to build something that allows for vibe-coding at full speed without compromising security. Traditional security tools aren’t effective, and manual pen-testing can’t keep up with the rapidly growing AI code This tool runs your code dynamically, finds vulnerabilities, and validates them through actual exploitation. You can either run it against your codebase or enter your (or someone…

    2025 · github.com

  18. 18
    VELA74

    Securely execute AI-generated & untrusted code

    Jun 2026 · vela-secure.vercel.app

  19. 19AT

    Hi Hacker News! We're launching Zalor, an agent testing platform. Agents often break when you tweak system prompts, swap models, or add tools. Zalor automatically generates test scenarios and evaluates your agent so you know it's reliable before deploying to production. We currently support the OpenAI Agents SDK and are onboarding other frameworks. A GitHub integration is coming so you can get feedback on every update. Looking forward to hearing feedback from people building agents.

    Mar 2026 · agents.zalor.ai

  20. 20

    Deterministic offline release evidence for AI agents

    Jul 2026 · iisacc-justmoong.github.io

  21. 21AE

    I've been working on a site [1] to give people control of their LLM workflows through AI evals - automated checks that, once defined, let you move fast without regressions and cut through hype with proof. That one-liner is aimed at software engineers, but I've spent my career helping cross-functional teams collaborate, and that's really what this is about. AI agents make powerful workflows very plausible, but only if teams can grow them incrementally without losing control - no vendor lock-in, no discipline silos, no blind trust in outputs. The site tries to meet different audiences where…

    Feb 2026 · ai-evals.io

  22. 22RB

    We built HALO (Hierarchal Agent Loop Optimizer), an open-source tool for debugging and optimizing AI agents using their execution traces. It’s a loop. Run your agent, feed the traces to HALO, get the report, apply the fixes, then re-run your agent. HALO takes in OTEL compliant traces from AI agents using tracing frameworks such as Langfuse, Arize/OpenInference, or even just plain JSONL. It uses an RLM (Recursive Language Model) to more efficiently break trace analysis into smaller subproblems in order to find recurring patterns across large amounts of data and fix systemic issues that…

    Jun 2026 · github.com

  23. 23

    The TLS for autonomous agent state.

    Jul 2026 · memora.optitransfer.ch

  24. 24

    Snapshot-test AI behavior in CI

    Jul 2026 · evalcore.cc

Ranked by how close each launch is in meaning, then by votes. Refine with a description →