pragma.vision Technology observatory

Verification register AI & Agents

Living definition

What is OpenAI SWE-bench Verified?

AI Coding & Software Agents AI & Agents

As of
2026-07-27
Revision
2026-07-28.1
Method
v1.3.0

Definition

The term, in context

AI-assisted draft · approved dataset

SWE-bench is a benchmark that evaluates how well AI coding agents can resolve real-world GitHub issues, and the human-audited subset of the dataset released by OpenAI filters out mislabeled or unsolvable problems, scoring agents by generating code patches that are checked against each repository's actual test suite.

Live readiness status

Status as of the current dateline

AI-assisted assembly · derived results

As of 2026-07-27, verified readiness is 26 (claimed 30, reported 29, gap 4) — Strong evidence strength; current signals suggest Proceed with caution.

The readiness fact belongs to the canonical Readiness Verdict for OpenAI SWE-bench Verified.

Dataset approval

Human editorial release

pragma.vision editorial — standing authorization (operator dev@soft.house, 2026-07-17)

Dataset
2026-07-28.1
Hash
sha256:02aa1ee2b9f805eacafe59552c1be4a62ac3e0edb8374314ff1adcfa3147c32f
Approved
2026-07-28

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.