nowfound

Alternatives

Products that do what Agent Compass does

Your AI Agent's Truth Graph to diagnose symptoms

  1. 1

    Trace, evaluate, and improve AI agents in production

    30d ago · telerik.com

  2. 2

    Evaluate AI workflows and reach 99% AI quality.

    Oct 2025

  3. 3
    Fabraix196

    Find gaps in your AI agents before users do

    May 2026

  4. 4
    AgentX151

    A reliable AI Agent to generate leads for your business

    2024

  5. 5

    See what breaks your AI agent and fix it automatically

    Jan 2026

  6. 6
    Polarity113

    The Self-Improvement Stack For agents

    May 2026

  7. 7

    Free go-to resource for all things AI agents automation

    Oct 2025

  8. 8

    Validate every PR with AI that runs tests for you

    Apr 2026

  9. 9
    Spiral161

    Analyze your reviews & support data with AI

    2025

  10. 10

    Quality control for your software factory

    Mar 2026

  11. 11
    Retrace101

    Debug AI agents by replaying and forking runs

    Jul 2026 · retraceai.tech

  12. 12
    Gauge111

    Agent Led Growth: Get written into every customer's codebase

    19d ago · withgauge.com

  13. 13

    Production failures become regression tests for AI agents

    26d ago · tracely-ai.com

  14. 14
    Nexus61

    AI-managed code quality and workflows built for AI teams

    2025

  15. 15

    Scam-proof your AI agents

    3d ago · agentlooker.ai

  16. 16AR
  17. 17MC

    Hi HN, I’ve been building AI agents and copilots, and kept running into a frustrating problem: they don’t fail loudly, they forget things quietly. Users re-explain preferences, agents contradict earlier responses, and context resets without any clear visibility into why. I built Memograph CLI as a debugging tool to analyze conversation transcripts and show: - what the agent forgot - where continuity broke - contradictions and repeated context - estimated token waste due to re-prompting It works locally and supports plain text or JSON transcripts. Example: $ memograph Output: Cognitive Drift…

    Feb 2026

  18. 18SF

    Hi HN, Over the past two years I’ve built and debugged a fair number of production pipelines—mainly retrieval‑augmented generation stacks, agent frameworks, and multi‑step reasoning services. A pattern emerged: most incidents weren’t outright crashes, but silent structural faults that slowly compromised relevance, accuracy, or stability. I began logging every recurring fault in a shared notebook. Colleagues started using the list for post‑mortems, so I turned it into a small public reference: 16 distinct failure modes (semantic drift after chunking, embedding/meaning mismatches,…

    2025 · github.com

  19. 197D

    hi all. i’ve been shipping a small open project that tries to answer that question with evidence, not vibes. in 70 days it reached \~800 stars. the core claim is simple: many AI failures are not noise. they repeat because the geometry and ordering underneath are stable. if so, we should be able to name each failure mode, set acceptance targets, and stop shipping the same bug twice. ### what it is * a compact Problem Map of 16 reproducible failure modes in RAG and agents. * each item has a minimal fix and measurable gates. examples: * Semantic ≠ Embedding: metric and normalization mismatch.…

    2025 · github.com

  20. 20AB

    Hi everyone! My team and I just open-sourced a bunch of cool agent dev tools: Invariant Explorer to visually inspect and understand AI traces and a testing framework, building on pytest.

    2024 · github.com

  21. 21AR

    If you're interested in exploring what LLM-based agent systems these days actually do to solve certain benchmarks such as SWEBench or WebArena, we created a small leaderboard with our team, that allows to view a lot of public and OSS agent results including all the runtime traces (the step-by-step reasoning behind the scenes). Looking at traces is actually quite interesting, as they reveal a lot about the inner working and shortcomings of current agent system, e.g. see https://explorer.invariantlabs.ai/u/invariant/webarena--SteP... for an example trace.

    2024 · explorer.invariantlabs.ai

  22. 22AA
  23. 23
    Currai1

    Find and fix failures in your AI agent conversations.

    13d ago · currai.app

  24. 24AT

    Hi Hacker News! We're launching Zalor, an agent testing platform. Agents often break when you tweak system prompts, swap models, or add tools. Zalor automatically generates test scenarios and evaluates your agent so you know it's reliable before deploying to production. We currently support the OpenAI Agents SDK and are onboarding other frameworks. A GitHub integration is coming so you can get feedback on every update. Looking forward to hearing feedback from people building agents.

    Mar 2026 · agents.zalor.ai

Ranked by how close each launch is in meaning, then by votes. Refine with a description →