pragma.vision Technology observatory & strategic foresight

Verification register AI & Agents

Comparative expertise

Autonomous AI coding agents

One fixed criteria frame. Every populated cell traces to one dated readiness verdict; missing evidence remains explicit.

Data as of
2026-08-20
Definition pinned
2026-07-17
Revision
2
Definition version
v1.0.0
Method
v1.3.0

Fixed matrix

Verdict fields, side by side

AI-assisted assembly · derived results

Scope: AI Coding & Software Agents · AI & Agents. No average, blend, composite, or estimate is produced.

Comparison of OpenAI Codex (GPT-5.5), AWS Kiro, Devin Desktop
Criterion OpenAI Codex (GPT-5.5) openai-codex-gpt-5-5 AWS Kiro aws-kiro Devin Desktop devin-desktop
Verified readiness verdict.readiness.verified 78 67 64
Hype gap verdict.readiness.gap 12 23 16
Evidence strength verdict.readiness.strength Strong Strong Critical
Recommended stance verdict.decision.answer Adopt with guardrails Wait for stronger evidence Proceed with caution
Latest evidence verdict.evidence_summary.latest_observed_on 2026-05-26 2026-06-01 2026-06-19

Human editorial

Synthesis and caveats

Human editorial · 2026-07-17

Synthesis

All three compared agents are generally available products: OpenAI Codex (the corpus row tracks the GPT-5.5 era; OpenAI released the GPT-5.6 family across Codex on 2026-07-09), AWS Kiro (GA since 2025-11-17, with team features and CLI support), and Devin Desktop (GA 2026-06-02 as the next generation of Windsurf, adding an agent command center with Agent Client Protocol support). Independent head-to-head evidence remains thin: the RuBench study (July 2026; 25 repository tasks, three runs each) measured Codex CLI with GPT-5.5 at 66.7% pass@1 against a best-measured 78.7%, but reports that most gaps are not statistically resolvable at that sample size. Practitioner and analyst reports converge on adoption being decided less by raw capability than by governance: output verification, nondeterminism, data confidentiality, enterprise support, and unpredictable token-based cost. The claimed-versus-verified gaps in the matrix (12–23 points) are consistent with that reading — vendor capability claims currently outpace independently verified production evidence for all three agents.

Caveats

Head-to-head evidence is one small benchmark plus review-scale reports, several from vendor or self-published sources (tier 3). One widely-cited Kiro reliability incident (April 2026) is attributed by AWS to misconfigured access controls rather than the agent itself. The Codex row's readiness reading predates the GPT-5.6-era refresh.

Approved by pragma.vision editorial — operator dev@soft.house · 2026-07-17

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.