nowfound

Alternatives

Products that do what AgentReplay does

Trace, Eval, compare, and improve AI agents locally

  1. 1
    AgentX523

    Evaluate AI agent, pinpoint issues, and fix with one click.

    Jun 2026 · agentx.so

  2. 2

    Trace, evaluate, and improve AI agents in production

    Aug 2026 · telerik.com

  3. 3
    Agenta362

    Open-source prompt management & evals for AI teams

    Nov 2025

  4. 4
    Retrace101

    Debug AI agents by replaying and forking runs

    Jul 2026 · retraceai.tech

  5. 5

    Test your AI apps with thousands of digital humans

    2025

  6. 6

    Open source, free, local debugger for AI agents

    May 2026 · raindrop.ai

  7. 7

    An open benchmark for AI agents that test APIs

    May 2026 · resources.kusho.ai

  8. 8

    Improve your LLM apps with open-source observability tool

    2024

  9. 9AS
  10. 10

    Visual debugging, tracing, and replay for agent workflows

    Apr 2026 · agenticlens.in

  11. 11OA

    Hi HN, we're Kiran and Vijay! Over the past two years, we have built a columnar storage engine for observability: logs, metrics, and traces. Today, it's exciting for us to show what we've built on top of that foundation: LLM Agent Observability. Given how non-deterministic agents are, storing all traces without sampling was critical for us. But these traces tend to be in the MBs, sometimes GBs - we needed to store them inexpensively. We also needed the queries and analyses to be fast. To meet both these goals, we store them in S3 in our own parquet-like file format, and query them using AWS…

    Jul 2026 · oodle.ai

  12. 12WE

    Hey HN! We’ve been building an MCP server to help AI-assisted web app developers by using browser agents to test whether changes made by an AI inside an editor actually work. We've been testing it on scenarios like verifying new flows in a UI, or checking that sending a chat request triggers a response. The idea is to let your coding agent both code and evaluate if what it did was correct. Here’s a short demo with Cursor: https://www.youtube.com/watch?v=_AoQK-bwR0w When building apps, we found the hardest part of AI-assisted coding isn’t the coding—it’s tedious point-and-click…

    2025 · github.com

  13. 13
    Foglamp101

    Ship AI agents you can actually see

    Jun 2026 · foglamp.dev

  14. 14RT
  15. 152C

    Single-agent LLMs suck at long-running complex tasks. We’ve open-sourced a multi-agent orchestrator that we’ve been using to handle long-running LLM tasks. We found that single LLM agents tend to stall, loop, or generate non-compiling code, so we built a harness for agents to coordinate over shared context while work is in progress. How it works: 1. Orchestrator agent that manages task decomposition 2. Sub-agents for parallel work 3. Subscriptions to task state and progress 4. Real-time sharing of intermediate discoveries between agents We tested this on a Putnam-level math problem, but the…

    Feb 2026 · github.com

  16. 16

    Deterministic offline release evidence for AI agents

    Jul 2026 · iisacc-justmoong.github.io

  17. 17AB

    Hi everyone! My team and I just open-sourced a bunch of cool agent dev tools: Invariant Explorer to visually inspect and understand AI traces and a testing framework, building on pytest.

    2024 · github.com

  18. 18WI

    At Laminar (https://github.com/lmnr-ai/lmnr) we're building open source AI observability platform in Rust. We obsess over instrumentation DX for our Python and TS SDKs and in this new blog we outline how we made the most seamless way of instrumenting recently released claude agent sdk

    Dec 2025 · laminar.sh

  19. 19
  20. 20

    Full observability for AI agents. Zero code changes.

    Apr 2026 · agentlens.techmatbd.com

  21. 21RB

    We built HALO (Hierarchal Agent Loop Optimizer), an open-source tool for debugging and optimizing AI agents using their execution traces. It’s a loop. Run your agent, feed the traces to HALO, get the report, apply the fixes, then re-run your agent. HALO takes in OTEL compliant traces from AI agents using tracing frameworks such as Langfuse, Arize/OpenInference, or even just plain JSONL. It uses an RLM (Recursive Language Model) to more efficiently break trace analysis into smaller subproblems in order to find recurring patterns across large amounts of data and fix systemic issues that…

    Jun 2026 · github.com

  22. 22MA

    We built meta-agent: an open-source library that automatically and continuously improves agent harnesses from production traces. Point it at an existing agent, a stream of unlabeled production traces, and a small labeled holdout set. An LLM judge scores unlabeled production traces as they stream. A proposer reads failed traces and writes one targeted harness update at a time, such as changes to prompts, hooks, tools, or subagents. The update is kept only if it improves holdout accuracy. On tau-bench v3 airline, meta-agent improved holdout accuracy from 67% to 87%. We open-sourced meta-agent.…

    Apr 2026 · github.com

  23. 23
    Farol2

    AI agent observability. One decorator, that's it.

    Apr 2026 · usefarol.dev

  24. 24

    Stop Testing AI Agents Manually Ship AI agents confidently.

    Jan 2026 · overseex.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →