pragma.vision Technology observatory & strategic foresight

Verification register AI & Agents

Readiness verdict

xAI Grok 4.6

A dated reading of what is claimed, reported, and independently verified in the current evidence.

As of
2026-08-20
Revision
1
Method
v1.3.0

Current reading

Readiness band and full integer triple

AI-assisted assembly · derived results

Readiness band

Ready

Primary summary from verified readiness

Confidence
58% · fresh
Computed at
2026-08-20T10:56:05.569988+00:00
Claimed
90

Public ambition and stated capability

Reported
84

Observed practitioner reporting

Verified
74

Independently supported evidence

Gap
+16

Claimed leads verified

Evidence strength Critical

Decision

What the current evidence supports

Human editorial judgment · 2026-08-20

Proceed with caution

Why
Matches GPT-5.6 Sol on the AA Index (61) at a stated $2/$6 per-million-token price — a genuine cost signal — but xAI's own published numbers show it trailing on the two benchmarks closest to actual agentic coding work (Terminal-Bench 26% vs ~34%, DeepSWE 65.9% vs 73%), and every comparative figure is self-reported five days post-launch.
Next
Run Grok 4.6 against Pragma.Vision's real agentic-coding workloads (Codex/Claude routing comparison) before adding it as a default model tier; re-check once independent (non-xAI-reported) benchmark verification exists.

Constraints

Blockers

No named blocker is present in the current public projection.

Evidence summary

Derived counts

AI-assisted assembly

Total
5
Tier 1
0
Tier 2
3
Tier 3
2
Supports
3
Contradicts
2
Context
0
Latest observed
2026-08-17

Counts and dates only. Raw signals, private excerpts, trust records, and internal corpus material are not published here.

Publication record

Revisions

Initial public reading

  1. 2026-08-17 Reading moved from ready to ready.

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.