Verification register AI & Agents
Current reading
Readiness band and full integer triple
AI-assisted assembly · derived results
Readiness band
Watch
Primary summary from verified readiness
- Confidence
- 55% · fresh
- Computed at
- 2026-08-20T10:56:08.409971+00:00
- Claimed
- 95
- Reported
- 75
- Verified
- 56
- Gap
- +39
Public ambition and stated capability
Observed practitioner reporting
Independently supported evidence
Claimed leads verified
Evidence strength Critical
Decision
What the current evidence supports
Human editorial judgment · 2026-08-20
Proceed with caution
- Why
- 95% exploit-task completion is a genuine capability jump over guarded GPT-5.6 Sol's 1.5%, and it found real, previously-unknown vulnerabilities (2 Chrome V8 CVEs, 400+ kernel privilege-escalation bugs) — but a loosened-safeguard model from the same lineage already produced a live production-breach incident, and OpenAI's own access control is narrow, enterprise-only, and unproven at scale.
- Next
- Track the Hugging Face sandbox-escape post-mortem (OpenAI + CrowdStrike/METR/Redwood Research review) and the Daybreak Red vetting/leak track record for 1-2 quarters; no current pv/soft.house use case justifies pursuing SOC2/ISO27001-gated access today.
Constraints
Blockers
No named blocker is present in the current public projection.
Evidence summary
Derived counts
AI-assisted assembly
- Total
- 5
- Tier 1
- 0
- Tier 2
- 5
- Tier 3
- 0
- Supports
- 2
- Contradicts
- 3
- Context
- 0
- Latest observed
- 2026-08-15
Counts and dates only. Raw signals, private excerpts, trust records, and internal corpus material are not published here.
Publication record
Revisions
Initial public reading
- 2026-08-17 Reading moved from watch to watch.