Comparative expertise
Autonomous AI coding agents
One fixed criteria frame. Every populated cell traces to one dated readiness verdict; missing evidence remains explicit.
- Data as of
- 2026-08-20
- Definition pinned
- 2026-07-17
- Revision
- 2
- Definition version
- v1.0.0
- Method
- v1.3.0
Fixed matrix
Verdict fields, side by side
AI-assisted assembly · derived results
Scope: AI Coding & Software Agents · AI & Agents. No average, blend, composite, or estimate is produced.
| Criterion | OpenAI Codex (GPT-5.5) openai-codex-gpt-5-5 | AWS Kiro aws-kiro | Devin Desktop devin-desktop |
|---|---|---|---|
| Verified readiness verdict.readiness.verified | 78 | 67 | 64 |
| Hype gap verdict.readiness.gap | 12 | 23 | 16 |
| Evidence strength verdict.readiness.strength | Strong | Strong | Critical |
| Recommended stance verdict.decision.answer | Adopt with guardrails | Wait for stronger evidence | Proceed with caution |
| Latest evidence verdict.evidence_summary.latest_observed_on | 2026-05-26 | 2026-06-01 | 2026-06-19 |
Human editorial
Synthesis and caveats
Human editorial · 2026-07-17
Synthesis
All three compared agents are generally available products: OpenAI Codex (the corpus row tracks the GPT-5.5 era; OpenAI released the GPT-5.6 family across Codex on 2026-07-09), AWS Kiro (GA since 2025-11-17, with team features and CLI support), and Devin Desktop (GA 2026-06-02 as the next generation of Windsurf, adding an agent command center with Agent Client Protocol support). Independent head-to-head evidence remains thin: the RuBench study (July 2026; 25 repository tasks, three runs each) measured Codex CLI with GPT-5.5 at 66.7% pass@1 against a best-measured 78.7%, but reports that most gaps are not statistically resolvable at that sample size. Practitioner and analyst reports converge on adoption being decided less by raw capability than by governance: output verification, nondeterminism, data confidentiality, enterprise support, and unpredictable token-based cost. The claimed-versus-verified gaps in the matrix (12–23 points) are consistent with that reading — vendor capability claims currently outpace independently verified production evidence for all three agents.
Caveats
Head-to-head evidence is one small benchmark plus review-scale reports, several from vendor or self-published sources (tier 3). One widely-cited Kiro reliability incident (April 2026) is attributed by AWS to misconfigured access controls rather than the agent itself. The Codex row's readiness reading predates the GPT-5.6-era refresh.
Approved by pragma.vision editorial — operator dev@soft.house · 2026-07-17