Alternatives
Products that do what AEVS does
proof-of-execution for AI agents
- 1

Trace, evaluate, and improve AI agents in production
30d ago · telerik.com
- 2

- 3

- 4

- 5

- 6

- 7Aegisora▲93
The narrow control plane for AI agent tool and API calls.
Aug 2026 · aegisora-ai.vercel.app
- 8

- 9

- 10

- 11
- 12

Proof of what your AI agent did, redacted by default
23d ago · aer.ktlsr.com
- 13

- 14

An open protocol for AI-agent activity events and control
5d ago · agenteventprotocol.io
- 15GV
Apr 2026 · github.com
- 16KP
AI agents increasingly execute real system actions: issuing refunds, modifying databases, deploying infrastructure, calling external APIs. Because agents retry steps, re-plan tasks, and run asynchronously, the same action can sometimes execute more than once. In production systems this can cause duplicate payouts, repeated mutations, or inconsistent state. Kybernis is a reliability layer that sits at the execution boundary of agent systems. When an agent calls a tool: 1. execution intent is captured 2. the action is recorded in an execution ledger 3. idempotency guarantees are attached 4.…
Mar 2026 · kybernis.io
- 17AB
Hey HN, I've been building Aether, a background agent that takes production errors from Sentry and attempts to turn them into verified pull requests. When a new error hits your Sentry project: 1. Sentry webhook fires with the stack trace, breadcrumbs, and context 2. Aether spins up an isolated Fly.io VM and clones the repo at the relevant commit 3. Agent analyzes the stack trace, reproduces the issue, proposes a fix 4. Starts the dev server, re-runs tests, and can verify the running app with Playwright (headless Chromium is pre-installed in every VM) 5. A review pass evaluates the diff…
Feb 2026
- 18

- 19AB
Hi everyone! My team and I just open-sourced a bunch of cool agent dev tools: Invariant Explorer to visually inspect and understand AI traces and a testing framework, building on pytest.
2024 · github.com
- 20WI
At Laminar (https://github.com/lmnr-ai/lmnr) we're building open source AI observability platform in Rust. We obsess over instrumentation DX for our Python and TS SDKs and in this new blog we outline how we made the most seamless way of instrumenting recently released claude agent sdk
Dec 2025 · laminar.sh
- 21AR
If you're interested in exploring what LLM-based agent systems these days actually do to solve certain benchmarks such as SWEBench or WebArena, we created a small leaderboard with our team, that allows to view a lot of public and OSS agent results including all the runtime traces (the step-by-step reasoning behind the scenes). Looking at traces is actually quite interesting, as they reveal a lot about the inner working and shortcomings of current agent system, e.g. see https://explorer.invariantlabs.ai/u/invariant/webarena--SteP... for an example trace.
2024 · explorer.invariantlabs.ai
- 22SA
Heya HN, excited to show off what I've been privately calling an AI cybersecurity tool built by AI skeptics. Two years ago we started a series of experiments with this philosophy of identifying small pieces of cognitive work where a human can very clearly map out the input data they need and the algorithm they'd follow to make a decision. This idea came partly out of frustration with the zeitgeist involving throwing broad AI features (e.g. useless chat bots) into products that end up unreliable and are targeting no clear problem a user might actually have. It feels kind of like a machete vs.…
2025 · semgrep.dev
- 23RS
What relai-sdk is an open-source toolkit for making AI agents reliable via a complete learning loop: simulate → evaluate → optimize. Why Agent runs are stochastic; tool-calls fail; hard to reproduce, measure, and fix at scale. It’s also hard to align behavior with goals across output quality/format, cost, and latency. We need a loop that integrates user feedback and LLM evaluators directly into the agent code (prompts, configs, models, graphs) without overfitting. How - Simulation: LLM personas, mocked MCP servers/tools, synthetic data; can condition on real traces - Evaluation:…
Oct 2025 · github.com
- 24

See — and block — what your AI agents actually do
5d ago · agentrec.io
Ranked by how close each launch is in meaning, then by votes. Refine with a description →