nowfound

Alternatives

Products that do what Trails does

Automated insights from LLM agent runs

  1. 1
    Retrace101

    Debug AI agents by replaying and forking runs

    Jul 2026

  2. 2
    Gauge235

    Your marketing agent for organic, paid, and AI search

    Mar 2026

  3. 3

    First AI agent for deep website monitoring

    2025

  4. 4
    Gauge111

    Agent Led Growth: Get written into every customer's codebase

    19d ago · withgauge.com

  5. 5

    Open-source LLM tracing for agent visibility

    Mar 2026

  6. 6

    Server-side analytics for AI & bot traffic

    Feb 2026

  7. 7

    AI agents to create, analyze, and automate surveys

    2025

  8. 8

    Track the bots and AI agents crawling your website

    May 2026

  9. 9
    Ara143

    Agentic Wispr flow computer-use-agent living in your notch

    May 2026

  10. 10

    Improve your LLM apps with open-source observability tool

    2024

  11. 11

    Your AI agent for automating browsers

    2025

  12. 12
    Glimpse106

    The competitive intelligence agent

    Jul 2026

  13. 13

    Your AI agents team, terminals, notes: one infinite canvas

    Jul 2026

  14. 14
    dolv47

    Your AI operator for content, CRM, and GTM execution

    27d ago · dolv.work

  15. 15LA

    We combined Stanford's ACE (agents learning from execution feedback) with the Reflective Language Model pattern. Instead of reading traces in a single pass, an LLM writes and runs Python in a sandbox to programmatically explore them - finding cross-trace patterns that single-pass analysis misses. The framework achieved 2x consistency improvement on τ2-bench.

    Mar 2026 · github.com

  16. 16AR

    If you're interested in exploring what LLM-based agent systems these days actually do to solve certain benchmarks such as SWEBench or WebArena, we created a small leaderboard with our team, that allows to view a lot of public and OSS agent results including all the runtime traces (the step-by-step reasoning behind the scenes). Looking at traces is actually quite interesting, as they reveal a lot about the inner working and shortcomings of current agent system, e.g. see https://explorer.invariantlabs.ai/u/invariant/webarena--SteP... for an example trace.

    2024 · explorer.invariantlabs.ai

  17. 17HL

    At testup.io we have been working for a while to bring artificial intelligence to the field of test automation. Just a few years ago, the primary challenge laid in accurately identifying UI elements following minor structural changes, such as updates to IDs or paths. The emergence of Large Language Models (LLMs) raised the bar for what it meant to be smart. Now, we anticipate the robot to do lots of things autonomously, such as retry in cases of unresponsiveness or handle minor error reports. A more challenging, but soon expected feature, would involve the test robot navigating your web shop…

    2024 · github.com

  18. 18IM

    Every time I wanted to use LLMs in my existing pipelines the integration was very bloated, complex, and too slow. This is why I created a lightweight library that works just like scikit-learn, the flow generally follows a pipeline-like structure where you “fit” (learn) a skill from sample data or an instruction set, then “predict” (apply the skill) to new data, returning structured results. High-Level Concept Flow Your Data --> Load Skill / Learn Skill --> Create Tasks --> Run Tasks --> Structured Results --> Downstream Steps And the bast part: Every step can be saved and reused as…

    2025 · github.com

  19. 19AP

    Hi HN, I’m a solo developer and built AgentWatch to solve a problem I kept running into while building AI agents: preventing runaway loops and unexpected LLM spend before requests reach the model. AgentWatch sits in front of OpenAI, Anthropic, Gemini, Bedrock, Azure OpenAI, Groq, and others to enforce budgets and runtime policies. I’d really appreciate your feedback. If you’re building AI agents, does this solve a problem you’ve experienced? I’d also love to hear what you’d improve or challenge.

    Jun 2026 · agent-watch.dev

  20. 20SE

    Hey HN! I built self-driving sim and eval at Waymo. Now I’m building Scorecard to bring that approach to agent eval: reproducible, automated scoring for AI. Scorecard lets you: - Run LLM-as-judge evals on agent workflows: test tool usage, multi-step reasoning, and task completion in CI/CD or in a playground. - Debug failures with OpenTelemetry traces: see which tool failed, why your agent looped, and where reasoning went wrong. - Collaborate on datasets, simulated agents, and evaluation metrics. Try it out → https://app.scorecard.io (free tier, no payment required!) Docs →…

    Oct 2025 · docs.scorecard.io

  21. 21

    Let's find your trading thesis

    24d ago

  22. 22GL

    I wanted to do a complete audit of my AWS account but was dissatisfied with the existing tools. Many of them are clunky to use, and their verbose scan outputs are difficult to understand. So, I built my own open-source tool that uses LLMs to summarize the scan results.

    2024 · guard.dev

  23. 23AA

    Looking for feedback on how Props can make your life easier as an LLM application developer.

    2024 · wwww.getprops.ai

  24. 24

    Can AI crawlers read your site?

    25d ago · querytrace.co

Ranked by how close each launch is in meaning, then by votes. Refine with a description →