Verification register AI & Agents
Current reading
Readiness band and full integer triple
AI-assisted assembly · derived results
Readiness band
Emerging
Primary summary from verified readiness
- Confidence
- 68% · fresh
- Computed at
- 2026-07-27T18:59:11.762519+00:00
- Claimed
- 30
- Reported
- 29
- Verified
- 26
- Gap
- +4
Public ambition and stated capability
Observed practitioner reporting
Independently supported evidence
Claimed leads verified
Evidence strength Strong
Decision
What the current evidence supports
Human editorial judgment · 2026-07-27
Proceed with caution
- Why
- Its own co-creator (OpenAI) publicly stopped reporting SWE-bench Verified for frontier models in 2026 after finding contamination and saturation — any register note that still treats a bare Verified score as proof of coding capability would repeat the exact failure Rule #11 (status reflects reality) guards against.
- Next
- Cite SWE-bench Verified scores only as a floor-level sanity check alongside a decontaminated benchmark (SWE-bench Pro or SWE-rebench); never as sole evidence of frontier coding capability in platform/model-selection write-ups. Note: the SWE-bench Pro pairing is our own editorial recommendation, not a claim OpenAI itself made.
Constraints
Blockers
No named blocker is present in the current public projection.
Evidence summary
Derived counts
AI-assisted assembly
- Total
- 6
- Tier 1
- 1
- Tier 2
- 4
- Tier 3
- 1
- Supports
- 2
- Contradicts
- 2
- Context
- 2
- Latest observed
- 2026-07-27
Counts and dates only. Raw signals, private excerpts, trust records, and internal corpus material are not published here.
Publication record
Revisions
Initial public reading
- 2026-07-27 Reading moved from emerging to emerging.