pragma.vision Technology observatory & strategic foresight

Verification register AI & Agents

Living definition

What is OpenAI SWE-bench Verified?

AI Coding & Software Agents AI & Agents

As of
2026-08-20
Revision
2026-08-17.1
Method
v1.3.0

Definition

The term, in context

AI-assisted draft · approved dataset

SWE-bench is a benchmark that evaluates how well AI coding agents can resolve real-world GitHub issues, and the human-audited subset of the dataset released by OpenAI filters out mislabeled or unsolvable problems, scoring agents by generating code patches that are checked against each repository's actual test suite.

Live readiness status

Status as of the current dateline

AI-assisted assembly · derived results

As of 2026-08-20, verified readiness is 26 (claimed 30, reported 29, gap 4) — Strong evidence strength; current signals suggest Proceed with caution.

The readiness fact belongs to the canonical Readiness Verdict for OpenAI SWE-bench Verified.

Dataset approval

Human editorial release

pragma.vision editorial — standing authorization (operator dev@soft.house, 2026-07-17)

Dataset
2026-08-17.1
Hash
sha256:0c37108a182f510fc8cf4d1dfc3a30339f2b59e62654201edea2f530a268707e
Approved
2026-08-17

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.