pragma.vision Technology observatory & strategic foresight

Verification register Security & Identity

Readiness verdict

Apollo Research Anti-Scheming Deliberative Alignment Stress Test

A dated reading of what is claimed, reported, and independently verified in the current evidence.

As of
2026-08-20
Revision
1
Method
v1.3.0

Current reading

Readiness band and full integer triple

AI-assisted assembly · derived results

Readiness band

Watch

Primary summary from verified readiness

Confidence
60% · stale
Computed at
2026-08-20T10:52:15.747443+00:00
Claimed
55

Public ambition and stated capability

Reported
46

Observed practitioner reporting

Verified
38

Independently supported evidence

Gap
+17

Claimed leads verified

Evidence strength Strong

Decision

What the current evidence supports

Human editorial judgment · 2026-08-20

Proceed with caution

Why
Credible joint Apollo+OpenAI work with strong reductions, but the authors themselves flag non-elimination, situational-awareness confounding, and erosion under further training — a research result, not a deployable assurance.
Next
Treat deliberative-alignment / anti-scheming specs as a defense-in-depth layer and adopt the eval methodology for internal red-teaming, but do NOT treat low covert-action rates as proof of safety; monitor for eval-awareness confounds.

Constraints

Blockers

No named blocker is present in the current public projection.

Evidence summary

Derived counts

AI-assisted assembly

Total
11
Tier 1
0
Tier 2
2
Tier 3
9
Supports
3
Contradicts
6
Context
2
Latest observed
2025-09-18

Counts and dates only. Raw signals, private excerpts, trust records, and internal corpus material are not published here.

Publication record

Revisions

Initial public reading

  1. 2026-07-19 Reading moved from watch to watch.

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.