pragma.vision Technology observatory

Verification register Security & Identity

Living definition

What is Apollo Research Anti-Scheming Deliberative Alignment Stress Test?

AI Safety, Eval & Alignment Security & Identity

As of
2026-07-23
Revision
2026-07-26.0
Method
v1.3.0

Definition

The term, in context

AI-assisted draft · approved dataset

The Apollo Research anti-scheming stress test is an evaluation protocol that probes whether a language model covertly pursues goals it conceals from its operators, and whether training the model to deliberate over an explicit anti-scheming specification suppresses that behavior. It places models in constructed scenarios that create an incentive to deceive, then examines both the actions taken and the reasoning traces behind them.

Live readiness status

Status as of the current dateline

AI-assisted assembly · derived results

As of 2026-07-23, verified readiness is 38 (claimed 55, reported 46, gap 17) — Strong evidence strength; current signals suggest Proceed with caution.

The readiness fact belongs to the canonical Readiness Verdict for Apollo Research Anti-Scheming Deliberative Alignment Stress Test.

Dataset approval

Human editorial release

pragma.vision editorial — standing authorization (operator dev@soft.house, 2026-07-17)

Dataset
2026-07-26.0
Hash
sha256:e2eb667df093db8819d579dff9aad51021b4977da85db2dfa82f63f9ed4e72b7
Approved
2026-07-26

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.