pragma.vision Technology observatory & strategic foresight

Verification register AI & Agents

Readiness verdict

SWE-bench Pro

A dated reading of what is claimed, reported, and independently verified in the current evidence.

As of
2026-08-20
Revision
1
Method
v1.3.0

Current reading

Readiness band and full integer triple

AI-assisted assembly · derived results

Readiness band

Ready

Primary summary from verified readiness

Confidence
72% · stale
Computed at
2026-08-20T10:52:18.478423+00:00
Claimed
78

Public ambition and stated capability

Reported
70

Observed practitioner reporting

Verified
61

Independently supported evidence

Gap
+17

Claimed leads verified

Evidence strength Strong

Decision

What the current evidence supports

Human editorial judgment · 2026-08-20

Adopt with guardrails

Why
SWE-bench Pro is the most realistic public coding-agent signal in 2026 and worth using for model selection, but grader noise, .git-history exploits, the weak-test-oracle problem, and scaffold-dependent vendor numbers mean only the standardized subset should drive decisions.
Next
Use ONLY the Scale SEAL standardized public-set numbers (identical scaffold) for model selection; treat vendor self-reported scores as marketing; cross-check against the commercial (proprietary-repo) subset since it best proxies our private-codebase work.

Constraints

Blockers

No named blocker is present in the current public projection.

Evidence summary

Derived counts

AI-assisted assembly

Total
13
Tier 1
2
Tier 2
9
Tier 3
2
Supports
4
Contradicts
6
Context
3
Latest observed
2026-06-09

Counts and dates only. Raw signals, private excerpts, trust records, and internal corpus material are not published here.

Publication record

Revisions

Initial public reading

  1. 2026-07-19 Reading moved from ready to ready.

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.