nowfound

Alternatives

Products that do what Reflect does

Self-Improving Layer Between Agent's Observability & Action

  1. 1

    Everything in OpenClaw's terminal, you can now do visually

    Apr 2026 · rectify.so

  2. 2FG

    Hi HN, I'm Antoine Zambelli, AI Director at Texas Instruments. I built Forge, an open-source reliability layer for self-hosted LLM tool-calling. What it does: - Adds domain-and-tool-agnostic guardrails (retry nudges, step enforcement, error recovery, VRAM-aware context management) to local models running on consumer hardware - Takes an 8B model from ~53% to ~99% on multi-step agentic workflows without changing the model - just the system around it - Ships with an eval harness and interactive dashboard so you can reproduce every number I wanted to run a handful of always-on agentic systems…

    May 2026 · github.com

  3. 3

    No-code automation without losing the human touch

    2021

  4. 4AS
  5. 5
    Clears376

    Move beyond AI coding to Agentic Software Delivery

    20d ago · clears.ai

  6. 6

    Trace, evaluate, and improve AI agents in production

    Aug 2026 · telerik.com

  7. 7
    Foil83

    An AI agent that monitors your AI agents

    Mar 2026

  8. 8
    oqoqo340

    Build evals and custom benchmarks for real-world tasks

    27d ago · oqoqo.ai

  9. 9
    Retrace101

    Debug AI agents by replaying and forking runs

    Jul 2026 · retraceai.tech

  10. 10
    Polarity113

    The Self-Improvement Stack For agents

    May 2026 · polarity.so

  11. 11

    An open source observability stack for your backends

    2023

  12. 12
    Reflect188

    Create automated tests without writing a line of code

    2020

  13. 13

    Real-time observability dashboard for OpenClaw AI agents

    Feb 2026

  14. 14

    Create beautiful, interactive charts without writing code

    2018

  15. 15
    Foglamp101

    Ship AI agents you can actually see

    Jun 2026 · foglamp.dev

  16. 16RT

    This project (Agents Observe) started as an exploration into building automation harnesses around claude code. I needed a way to see exactly what teams of agents were doing in realtime and to filter and search their output. A few interesting learnings from building and using this: - Claude code hooks are blocking - performance degrades rapidly if you have a lot of plugins that use hooks - Hooks provide a lot more useful info than OTEL data - Claude's jsonl files provide the full picture - Lifecycle management of MCP processes started by plugins is a bit kludgy at best The biggest takeaway is…

    Apr 2026 · github.com

  17. 17
    Tracea80

    Datadog for AI agents with traces, RCA, and team memory

    May 2026 · tracea.dev

  18. 18RB

    We built HALO (Hierarchal Agent Loop Optimizer), an open-source tool for debugging and optimizing AI agents using their execution traces. It’s a loop. Run your agent, feed the traces to HALO, get the report, apply the fixes, then re-run your agent. HALO takes in OTEL compliant traces from AI agents using tracing frameworks such as Langfuse, Arize/OpenInference, or even just plain JSONL. It uses an RLM (Recursive Language Model) to more efficiently break trace analysis into smaller subproblems in order to find recurring patterns across large amounts of data and fix systemic issues that…

    Jun 2026 · github.com

  19. 19MR

    The most common failures for production agents are behavioral: looping, reasoning leakage, user frustration, and more. Using a frontier model like GPT or Sonnet to judge every turn is too expensive and slow to run at scale. To solve this, we built Reflexes: semantic signals from agent traces, served fast and cheap over API. Built on custom kernels and a custom inference engine forked from vLLM. Under the hood, it is a small LLM architected around multi-head inference. Small models need to be trained for specific tasks, but running 50 separate small models on the same input for 50 tasks makes…

    Jun 2026

  20. 20RT
  21. 21WE

    Hey HN! We’ve been building an MCP server to help AI-assisted web app developers by using browser agents to test whether changes made by an AI inside an editor actually work. We've been testing it on scenarios like verifying new flows in a UI, or checking that sending a chat request triggers a response. The idea is to let your coding agent both code and evaluate if what it did was correct. Here’s a short demo with Cursor: https://www.youtube.com/watch?v=_AoQK-bwR0w When building apps, we found the hardest part of AI-assisted coding isn’t the coding—it’s tedious point-and-click…

    2025 · github.com

  22. 22HO

    I'm Josh, founder of Synth. We've been working on coding agent optimization with method like GEPA and MIPRO (the latter of which, I helped to originally develop), agent evaluation via methods like RLMs, and large scale deployment for training and inference. We've also worked on patterns for memory, processing live context, and managing agent actions, combining it all in a single stack called Horizons. With the release of OpenAI's Frontier and the consumer excitement around OpenClaw, we think the timing is right to release a v0. It integrates with our sdk for evaluation and optimization but…

    Feb 2026 · github.com

  23. 23

    Observability for agents

    Mar 2026

  24. 24

    Deterministic offline release evidence for AI agents

    Jul 2026 · iisacc-justmoong.github.io

Ranked by how close each launch is in meaning, then by votes. Refine with a description →