nowfound

Alternatives

Products that do what Parity: Auto-evals for harness changes does

Catch AI behavior changes before they ship

  1. 1

    Monitor brand & link visibility on ChatGPT, Perplexity & AIO

    2024

  2. 2

    How visible is your brand and content on AI searches?

    2024

  3. 3
    Polarity113

    The Self-Improvement Stack For agents

    May 2026 · polarity.so

  4. 4
    Fruitful424

    Track competitors instantly. Juicy insights every day.

    Oct 2025

  5. 5

    Your AI search ranking tool—for ChatGPT, Gemini, & Claude.

    2025

  6. 6

    Your AI agent for automating browsers

    2025

  7. 7
    Atla493

    Automatically detect errors in your AI agents

    Sep 2025

  8. 8

    Open-source unified interface for agent harnesses

    22d ago · harnessrouter.ai

  9. 9

    Validate every PR with AI that runs tests for you

    Apr 2026 · qa.tech

  10. 10

    Trace, evaluate, and improve AI agents in production

    Aug 2026 · telerik.com

  11. 11

    A coding agent that can refine its own harness

    28d ago · primeintellect.ai

  12. 12

    Your AI Agent's Truth Graph to diagnose symptoms

    Sep 2025

  13. 13
    Foil83

    An AI agent that monitors your AI agents

    Mar 2026 · getfoil.ai

  14. 14AS
  15. 15WB

    Hey HN, We’re two developers (co-founders) with a team of 20 who got tired of spending hours reviewing PRs, so we built Infinitcode.ai, an AI-powered code reviewer that: - *Summarizes PRs in plain English*: No more deciphering 1,000-line diff jungles - *Catches more than bugs*: Security holes, performance pitfalls, code smells, even typos (yes, we’ll flag “vurnerabilities” and vulnerabilities) - *Zero onboarding*: Works instantly—no “let me learn your codebase for weeks” nonsense. Why we’re posting: We’re in alpha and need brutal honesty. Roast our tool, mock our UI, or tell us why AI will…

    2025 · infinitcode.ai

  16. 16

    Increase your brand's AI visibility on autopilot

    Jul 2026 · promptscout.app

  17. 17WE

    Hey HN! We’ve been building an MCP server to help AI-assisted web app developers by using browser agents to test whether changes made by an AI inside an editor actually work. We've been testing it on scenarios like verifying new flows in a UI, or checking that sending a chat request triggers a response. The idea is to let your coding agent both code and evaluate if what it did was correct. Here’s a short demo with Cursor: https://www.youtube.com/watch?v=_AoQK-bwR0w When building apps, we found the hardest part of AI-assisted coding isn’t the coding—it’s tedious point-and-click…

    2025 · github.com

  18. 18AH
  19. 19RT

    This project (Agents Observe) started as an exploration into building automation harnesses around claude code. I needed a way to see exactly what teams of agents were doing in realtime and to filter and search their output. A few interesting learnings from building and using this: - Claude code hooks are blocking - performance degrades rapidly if you have a lot of plugins that use hooks - Hooks provide a lot more useful info than OTEL data - Claude's jsonl files provide the full picture - Lifecycle management of MCP processes started by plugins is a bit kludgy at best The biggest takeaway is…

    Apr 2026 · github.com

  20. 20
    Regent11

    Know when your AI changes behavior

    Apr 2026 · portal.regentai.in

  21. 21AB

    Hi there, HN! We’re Jai and Sanket from DeepSource (YC W20), and today we’re launching Autofix Bot, a hybrid static analysis + AI agent purpose-built for in-the-loop use with AI coding agents. AI coding agents have made code generation nearly free, and they’ve shifted the bottleneck to code review. Static-only analysis with a fixed set of checkers isn’t enough. LLM-only review has several limitations: non-deterministic across runs, low recall on security issues, expensive at scale, and a tendency to get ‘distracted’. We spent the last 6 years building a deterministic, static-analysis-only…

    Dec 2025

  22. 22MA

    We built meta-agent: an open-source library that automatically and continuously improves agent harnesses from production traces. Point it at an existing agent, a stream of unlabeled production traces, and a small labeled holdout set. An LLM judge scores unlabeled production traces as they stream. A proposer reads failed traces and writes one targeted harness update at a time, such as changes to prompts, hooks, tools, or subagents. The update is kept only if it improves holdout accuracy. On tau-bench v3 airline, meta-agent improved holdout accuracy from 67% to 87%. We open-sourced meta-agent.…

    Apr 2026 · github.com

  23. 23

    Research-grounded, harness-agnostic skills for AI coding agents - SteveVitali/agent-skills

    Aug 2026 · github.com

  24. 24OS

    GitHub - https://github.com/vostride/agent-qa Live Demos - https://vostride.com/demo/agent-qa

    May 2026 · vostride.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →