Benchmark: AI doesn't find bugs unless you tell it what's wrong
Given a real repository, an agent must discover & repair as many bugs as they can. Agents are not given any hint about the type of bug or its location.
Benchmark: AI doesn't find bugs unless you tell it what's wrong launched on October 1, 2026 with 5 votes, #38 of 279 launches that month and more than 75% of that year's launches.
What it does
Given a real repository, an agent must discover & repair as many bugs as they can. Agents are not given any hint about the type of bug or its location. Per-repository resource use includes non-deprecated retries and is averaged over evaluated repositories. Most existing software engineering benchmarks evaluate coding agents on concrete, well-specified tasks, commonly by providing a codebase together with a user-reported issue to resolve. However, as users delegate increasingly broad outcomes to coding agents, the natural next step is for agents to determine not only how to perform useful work, but also what useful work needs to be done. An agent entrusted with a repository should be able to…from swesweep.com
Does a similar job
all alternatives →
Detail, a Bug FinderDec 2025 · detail.dev · ▲67Hi HN, tl;dr we built a bug finder that's working really well, especially for app backends. Try it out and send us your thoughts! Long story below. -------------------------- We originally set out to work on technical debt. We had all seen codebases with a lot of debt, so we had personal grudges about the problem, and AI seemed to be making it a lot worse. Tech debt also seemed like a great problem for AI because: 1) a small portion of the work is thinky and strategic, and then the bulk of the execution is pretty mechanical, and 2) when you're solving technical debt, you're usually trying to…

TrailsJan 2026 · trails-red.vercel.app · ▲81Hey all, Justin here. I previously built Phind, the AI search engine for developers. One of the biggest problems we had there was figuring out what went wrong with bad searches. We had tons of searches per day, but less than 1% of users gave any explicit feedback. So we were either manually digging through searches or making general system improvements and hoping they helped. This problem gets harder with agents. Traces are longer and more complex. It takes more effort to review them, so I'm building a tool that lets you analyze LLM outputs directly to help developers of LLM apps and agents…
Git for AI AgentsMay 2026 · github.com · ▲129hi guys. been working on something i think is fundamentally missing in today's workflow with ai agents. vcs. i find myself struggling with questions that agents can't answer like "why did you do it?", "when did u delete this folder? why?", etc. or trying to /rewind (after a /compact...) or basically `bisect` to find when and why something was done by the agent in the current / previous session. just like git did for code, i think we are the same core capabilities with ai agents so... i developed an open source solution for that (currently supporting claude code) would love to…
Fix My Agent (FMA)Jan 2026 · docs.futureagi.com · ▲50See what breaks your AI agent and fix it automatically
More ai this month
the category →
E-ink bird frame for Raspberry Pi - real-time bird detection by audio, fully local AI, rendered as real, hand-cut 1800s bird illustrations. - arnegiacomo/fugleramme
AI · 17d ago · github.com




PixVerse R2▲381A real-time world model you can explore and change
AI · 10d ago · world.pixverse.video
Launched alongside, October 2026
the whole month →- VO
AI · 1d ago · stoppels.ch

Paintings by AI models: programs run through a simulation of oil paint.
AI · 21h ago · stillwet.art


Dev tools · 14h ago · github.com

