TL;DR AI can now build you an app in minutes, but the bugs in that app still get found the old way: by the people using it. SWEeper-Bench asks whether an agent can find and fix those bugs first. The benchmark is built from 200 real bugs in 200 open-source web apps, each one hidden behind a particular sequence of clicks. Instead of a bug report, the agent only gets a broad request such as find and fix issues in board sharing
, and has to discover the bug on its own. The best of 15 frontier agents fixes 59% of the bugs.
Who finds the bugs?
More and more, when we ask AI for something, it responds by writing software on the fly: a sign-up page for the team trip, a booking form, a little tool for the office. A few minutes later the app exists, and the people it’s for are already using it.
And they use it in all kinds of ways nobody planned for. Grandma opens it on an old iPad with the text turned way up, and the send button gets pushed off the screen. Your cousin opens it from Tokyo and sees the wedding listed a day early. Two friends press send at the exact same second, and one of their replies quietly vanishes. Someone on a train with a shaky connection taps send twice and ends up on the guest list twice.
What makes bugs like these so hard to catch is that you can’t see them in the code. Read it line by line and it looks fine, and the unit tests pass. The bug only shows up when a real person uses the app on a particular device, in a particular order, or at the same moment as someone else. To find it, you have to actually use the app.
So right now, users are the testers. They hit the bug, they complain, the complaint makes its way back to whoever built the thing, and everyone waits for a fix. It happens to the biggest companies in the world:
This Glitch Can FULLY CRASH the new iPhone 18 Pro
TechDroider (@techdroider), September 27, 2026
New iOS glitch has been discovered which can make the iPhone completely UNRESPONSIVE and mess up the entire UI. Apple NEEDS to take action immediately and push a HOTFIX update ASAP.
This is not a minor visual bug. Your iPhone pic.x.com/pFYqcnbnt9
With AI writing apps on the fly, this is about to happen a lot more often. Those apps are usually made for a handful of people, with no testing team anywhere in between, so the people using them end up as the only testers. What we want instead is for the agent that built the app to use it the way people will, find what’s broken, and fix it before anyone else notices.
That’s the part no benchmark measures yet. SWE-bench and its successors hand the agent a bug report someone already wrote, so they measure whether an agent can fix a bug that has already been found, not whether it can find the bug in the first place. SWEeper-Bench starts before anyone has reported anything.
The best agent fixes 59%
Each task gives an agent the app’s code, the running app, a browser, and a broad request that only names part of the app, such as Find and fix issues in Focalboard’s board sharing and membership.
We tried 15 frontier agents, and the best one, Grok 4.6 in Cursor, fixed 59%. Nine of the fifteen fixed fewer than half.
Target: the target behavior test passes, meaning the verifier repeats the interactions that expose the bug and sees the intended behavior. Preserv.: the preservation test passes, meaning nearby features still work. Pass requires both. All are percentages of the 200 tasks; time is mean ± standard deviation per task. Click a column to sort.
How it works
The agent receives the app’s codebase and an open-ended prompt like Find and fix issues in SearXNG’s search suggestion interface.
It can run any tests it likes in a sandbox with a browser installed. Once it produces a code patch, we run the patched app and have a verifier agent use it in a browser, following step-by-step instructions written for that task. Each task has two tests: a target test, which repeats the interactions that expose the bug and checks that it’s gone, and a preservation test, which checks that nearby features still work. A task passes only if both do.
How reliable is the verifier?
We ran both behavior tests on every buggy version and every reference fix, with three verifier setups, three times each. A perfect verifier fails the target test on the buggy version and passes everything else. All three setups come within one point of that and vary by at most 0.8 points between runs. In other words, the behavior tests are robust: they give the same verdicts no matter which model or harness runs them. We use the cheapest setup, GPT-5.6 Luna in Codex.
| Verifier | Buggy version | Reference fix | Cost / task | ||
|---|---|---|---|---|---|
| Target ↓ | Preserv. ↑ | Target ↑ | Preserv. ↑ | ||
| 0.2 ± 0.3 | 99.2 ± 0.8 | 99.8 ± 0.3 | 99.3 ± 0.3 | $0.16 | |
| 0.8 ± 0.3 | 99.8 ± 0.3 | 99.8 ± 0.3 | 99.7 ± 0.3 | $0.64 | |
| 0.2 ± 0.3 | 99.8 ± 0.3 | 99.8 ± 0.3 | 100.0 ± 0.0 | $3.84 | |
Every task is a real bug that someone fixed in a real open-source web app, with 200 different apps spread across seven kinds of software. These bugs are hard to spot. Many only appear after a long sequence of steps, sometimes across more than one account, and they often come from several parts of the app working together, like the interface, the server and the stored data, rather than from one obviously wrong line. Often nothing crashes at all: the app just quietly does the wrong thing, and only someone paying close attention would notice.
Test case example: shared team board doesn’t show up
Focalboard is an open-source project board, kind of like Trello. When you share a board with a teammate, it’s supposed to show up in their sidebar. In the buggy version it never does, and nothing looks wrong from your side, so the only way to notice is to share a board, log in as your teammate, and look. Here’s our verifier doing exactly that:
The fix turns out to be a single function call. The hard part is noticing there’s anything to fix: four agents made it to that empty sidebar, and only two of them thought it was a bug. Claude Opus 5 checked the unit tests, saw that they expect no sidebar update, and concluded the empty sidebar was intentional. The existing tests had the bug built into them, so they pointed the agent the wrong way.
Where do agents stumble?
To see where agents go wrong, we split each run into stages. About a quarter of the time, the agent never gets to the feature where the bug lives. Another quarter of the time it gets there and still misses the bug: it never does the thing that triggers it, or it triggers it and doesn’t recognize what it’s looking at. Once an agent actually sees the bug, diagnosing and fixing it almost never goes wrong.
Fixing, on the other hand, is the easy part: when we tell agents what the bug is, every agent we tried passes about 96% of the tasks.
More time doesn’t fix it
You might expect more time to help, but it mostly doesn’t. Grok 4.6 spends about 12 minutes per task, while Qwen 3.8 Flash spends 102 and scores a little lower. Giving the same agent a bigger budget only helps up to about 40 minutes, and then it flattens out.
Closing the loop
Today, the loop that makes interactive software work still runs through people: someone builds the app, the people using it find what breaks, and fixes follow. As AI builds more of the software we use, it can start running that loop itself, using the app the way a person would, noticing what’s off, and fixing it, again and again, until what reaches people just works. SWEeper-Bench measures how close agents are to doing that on their own.
Citation
@article{yao2026sweeperbench,
title = {SWEeper-Bench: Can Agents Discover Bugs in Interactive Software?},
author = {Yao, Yang and Chen, Haozhe and Kang, Bingyi and
Narasimhan, Karthik R and Liu, Zhuang},
journal = {arXiv preprint},
year = {2026}
}
Acknowledgments
We thank Modal for providing the cloud sandboxes in which we ran all agent and verifier experiments.