pragma.vision Technology observatory & strategic foresight

Verification register Security & Identity

Readiness verdict

Inspect / Inspect Evals (UK AISI)

A dated reading of what is claimed, reported, and independently verified in the current evidence.

As of
2026-08-20
Revision
1
Method
v1.3.0

Current reading

Readiness band and full integer triple

AI-assisted assembly · derived results

Readiness band

Mature

Primary summary from verified readiness

Confidence
73% · stale
Computed at
2026-08-20T10:53:43.035197+00:00
Claimed
85

Public ambition and stated capability

Reported
84

Observed practitioner reporting

Verified
76

Independently supported evidence

Gap
+9

Claimed leads verified

Evidence strength Critical

Decision

What the current evidence supports

Human editorial judgment · 2026-08-20

Adopt with guardrails

Why
Government-backed, MIT-licensed, very high activity (inspect_ai 2.2k stars / 6,273 commits), and the de-facto frontier-eval standard; the strongest production-grade choice in this discipline, with the caveat that community evals need per-eval validation.
Next
Stand up Inspect AI as the standard eval harness for model selection; pin Python 3.12 and a curated subset of evals, validating each chosen eval against its source paper before trusting scores.

Constraints

Blockers

No named blocker is present in the current public projection.

Evidence summary

Derived counts

AI-assisted assembly

Total
10
Tier 1
0
Tier 2
3
Tier 3
7
Supports
8
Contradicts
2
Context
0
Latest observed
2026-06-24

Counts and dates only. Raw signals, private excerpts, trust records, and internal corpus material are not published here.

Publication record

Revisions

Initial public reading

  1. 2026-07-19 Reading moved from mature to mature.

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.