The agent is given the issue text and a repository snapshot, then proposes a code change (patch). The change is applied and the repository's tests are run (including the bug-reproducing tests and regression tests). A task counts as resolved when all required tests pass; the headline metric is the percentage of resolved issues (pass@1) across the 500 tasks.
The full SWE-bench contains unsolvable or poorly specified tasks that depress and add noise to scores; Verified provides a reliable, human-confirmed subset for comparing coding agents.
The original benchmark for resolving GitHub issues (arXiv 2310.06770, ICLR 2024).
A human-validated 500-task subset; a collaboration with OpenAI Preparedness (August 13, 2024).