Living definition
What is OpenAI SWE-bench Verified?
AI Coding & Software Agents AI & Agents
- As of
- 2026-07-27
- Revision
- 2026-07-28.1
- Method
- v1.3.0
Definition
The term, in context
AI-assisted draft · approved dataset
SWE-bench is a benchmark that evaluates how well AI coding agents can resolve real-world GitHub issues, and the human-audited subset of the dataset released by OpenAI filters out mislabeled or unsolvable problems, scoring agents by generating code patches that are checked against each repository's actual test suite.
Live readiness status
Status as of the current dateline
AI-assisted assembly · derived results
As of 2026-07-27, verified readiness is 26 (claimed 30, reported 29, gap 4) — Strong evidence strength; current signals suggest Proceed with caution.
The readiness fact belongs to the canonical Readiness Verdict for OpenAI SWE-bench Verified.
Dataset approval
Human editorial release
pragma.vision editorial — standing authorization (operator dev@soft.house, 2026-07-17)
- Dataset
- 2026-07-28.1
- Hash
- sha256:02aa1ee2b9f805eacafe59552c1be4a62ac3e0edb8374314ff1adcfa3147c32f
- Approved
- 2026-07-28