pragma.vision Technology observatory & strategic foresight

Verification register AI & Agents

Readiness verdict

OpenAI Codex (GPT-5.5)

A dated reading of what is claimed, reported, and independently verified in the current evidence.

As of
2026-08-20
Revision
1
Method
v1.3.0

Current reading

Readiness band and full integer triple

AI-assisted assembly · derived results

Readiness band

Mature

Primary summary from verified readiness

Confidence
73% · stale
Computed at
2026-08-20T10:53:07.934357+00:00
Claimed
90

Public ambition and stated capability

Reported
85

Observed practitioner reporting

Verified
78

Independently supported evidence

Gap
+12

Claimed leads verified

Evidence strength Strong

Decision

What the current evidence supports

Human editorial judgment · 2026-08-20

Adopt with guardrails

Why
We already run Codex (gpt-5.5 xhigh) as the implementer in trio/autonomous-conductor; its agentic coding (Terminal-Bench 82.7%) and ~40% token efficiency are strong, but the false-completion/misalignment findings justify our existing diff-verification guardrails rather than blind trust.
Next
Keep Codex as the S10 implementer but enforce the existing review protocol (pre/post git snapshot + diff audit) and never let it self-attest completion; treat its 'done' claims as untrusted given the 29% false-completion finding

Constraints

Blockers

No named blocker is present in the current public projection.

Evidence summary

Derived counts

AI-assisted assembly

Total
12
Tier 1
0
Tier 2
9
Tier 3
3
Supports
4
Contradicts
4
Context
4
Latest observed
2026-05-26

Counts and dates only. Raw signals, private excerpts, trust records, and internal corpus material are not published here.

Publication record

Revisions

Initial public reading

  1. 2026-07-19 Reading moved from mature to mature.

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.