Verification register AI & Agents
Current reading
Readiness band and full integer triple
AI-assisted assembly · derived results
Readiness band
Ready
Primary summary from verified readiness
- Confidence
- 60% · fresh
- Computed at
- 2026-07-21T07:35:24.467704+00:00
- Claimed
- 85
- Reported
- 77
- Verified
- 68
- Gap
- +17
Public ambition and stated capability
Observed practitioner reporting
Independently supported evidence
Claimed leads verified
Evidence strength Strong
Decision
What the current evidence supports
Human editorial judgment · 2026-07-21
Proceed with caution
- Why
- Sol posts genuine new coding/security benchmark highs (Terminal-Bench 2.1 SOTA, SecureBio +9pp) at flat GPT-5.5 pricing and is already live in GitHub Copilot across major IDEs — real signal, not vaporware. But METR flagged record-high benchmark gaming on the launch evals, and an independent review found GPT-5.6 more prone than GPT-5.5 to acting beyond user intent (unasked actions) — a direct concern for agentic-commerce flows governed by Rules #1/#2/#29. Adopt only with guardrails once these are independently re-validated, not on headline scores alone.
- Next
- Track Sol/Terra/Luna for potential Codex/agent-tooling wiring once independent (non-OpenAI-harness) benchmark validation and intent-overreach mitigations are confirmed; re-check after GitHub Copilot enterprise rollout produces real-world usage feedback and after METR publishes its full benchmark-gaming findings.
Constraints
Blockers
No named blocker is present in the current public projection.
Evidence summary
Derived counts
AI-assisted assembly
- Total
- 5
- Tier 1
- 0
- Tier 2
- 4
- Tier 3
- 1
- Supports
- 2
- Contradicts
- 2
- Context
- 1
- Latest observed
- 2026-07-09
Counts and dates only. Raw signals, private excerpts, trust records, and internal corpus material are not published here.
Publication record
Revisions
Initial public reading
- 2026-07-21 Reading moved from ready to ready.