pragma.vision Technology observatory & strategic foresight

Verification register Security & Identity

Readiness verdict

METR RE-Bench

A dated reading of what is claimed, reported, and independently verified in the current evidence.

As of
2026-08-20
Revision
1
Method
v1.3.0

Current reading

Readiness band and full integer triple

AI-assisted assembly · derived results

Readiness band

Ready

Primary summary from verified readiness

Confidence
67% · stale
Computed at
2026-08-20T10:53:26.278449+00:00
Claimed
80

Public ambition and stated capability

Reported
75

Observed practitioner reporting

Verified
70

Independently supported evidence

Gap
+10

Claimed leads verified

Evidence strength Strong

Decision

What the current evidence supports

Human editorial judgment · 2026-08-20

Adopt with guardrails

Why
Most mature of the three — peer-reviewed (ICLR/OpenReview), MIT-licensed, actively maintained with continual frontier-model updates through 2026 and an established human baseline. Reward-hacking and saturation are real but manageable with manual transcript review.
Next
Pilot RE-Bench (plus HCAST/SWAA time-horizon tasks) as a capability tripwire in any internal frontier-model-routing or agent-autonomy gate; pair scores with reward-hacking transcript inspection rather than trusting normalized scores alone.

Constraints

Blockers

No named blocker is present in the current public projection.

Evidence summary

Derived counts

AI-assisted assembly

Total
10
Tier 1
0
Tier 2
0
Tier 3
10
Supports
3
Contradicts
3
Context
4
Latest observed
2026-05-08

Counts and dates only. Raw signals, private excerpts, trust records, and internal corpus material are not published here.

Publication record

Revisions

Initial public reading

  1. 2026-07-19 Reading moved from ready to ready.

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.