Verification register AI & Agents
Current reading
Readiness band and full integer triple
AI-assisted assembly · derived results
Readiness band
Ready
Primary summary from verified readiness
- Confidence
- 73% · stale
- Computed at
- 2026-08-20T10:52:24.885664+00:00
- Claimed
- 83
- Reported
- 75
- Verified
- 67
- Gap
- +16
Public ambition and stated capability
Observed practitioner reporting
Independently supported evidence
Claimed leads verified
Evidence strength Strong
Decision
What the current evidence supports
Human editorial judgment · 2026-08-20
Adopt with guardrails
- Why
- Open-weight Apache-2.0 27B at flagship-level coding scores on commodity hardware, plus an API tier that undercuts Claude Opus on cost, is a real cost/capability lever — but selective benchmark framing, unsettled Plus pricing, and tier/version sprawl all demand independent validation before trusting vendor leaderboard claims.
- Next
- Pilot the Apache-2.0 Qwen3.6-27B (77.2 SWE-bench Verified, ~18GB-class hardware) for self-hosted coding/tool-use against the current pipeline; A/B the Plus API on a non-PII agentic workload to confirm the 78.8 SWE-bench claim on our own evals, and pin the actual per-token price (preview vs paid) before any production swap.
Constraints
Blockers
No named blocker is present in the current public projection.
Evidence summary
Derived counts
AI-assisted assembly
- Total
- 13
- Tier 1
- 4
- Tier 2
- 6
- Tier 3
- 3
- Supports
- 3
- Contradicts
- 6
- Context
- 4
- Latest observed
- 2026-04-25
Counts and dates only. Raw signals, private excerpts, trust records, and internal corpus material are not published here.
Publication record
Revisions
Initial public reading
- 2026-07-19 Reading moved from ready to ready.