pragma.vision Technology observatory & strategic foresight

Verification register AI & Agents

Comparative expertise

Multimodal vision-language models

One fixed criteria frame. Every populated cell traces to one dated readiness verdict; missing evidence remains explicit.

Data as of
2026-08-05
Definition pinned
2026-08-05
Revision
1
Definition version
v1.0.0
Method
v1.3.0

Fixed matrix

Verdict fields, side by side

AI-assisted assembly · derived results

Scope: Multimodal & Vision-Language Models · AI & Agents. No average, blend, composite, or estimate is produced.

Comparison of Claude Opus 4.5 (Vision / Computer Use), Gemini 3 Pro (Vision), GLM-4.5V, Qwen3-VL
Criterion Claude Opus 4.5 (Vision / Computer Use) claude-opus-4-5-vision Gemini 3 Pro (Vision) gemini-3-pro-vision GLM-4.5V glm-4-5v Qwen3-VL qwen3-vl
Verified readiness verdict.readiness.verified 63 64 64 83
Hype gap verdict.readiness.gap 22 16 21 7
Evidence strength verdict.readiness.strength Strong Strong Strong Strong
Recommended stance verdict.decision.answer Adopt with guardrails Adopt with guardrails Proceed with caution Adopt with guardrails
Latest evidence verdict.evidence_summary.latest_observed_on 2026-05-28 2025-12-05 2025-08-12 2026-03-31

Human editorial

Synthesis and caveats

Human editorial · 2026-08-05

Synthesis

The four systems sit in materially different lifecycle states, which is itself the most load-bearing fact for anyone reading this cluster today. Google's Gemini 3 Pro Preview -- launched 2025-11-18 as an explicitly preview-labeled endpoint -- was shut down by Google on 2026-03-09; the API alias now resolves to Gemini 3.1 Pro Preview, and Google's own historical documentation (updated 2026-07-21) marks computer-use support as no longer available, contradicting its own 2026-01-29 changelog entry that had announced computer-use support for the same model. Anthropic's Claude Opus 4.5 (announced 2025-11-24, immediately available across the Claude apps, API and all three major cloud platforms) remains live and documented as of this reading, even though Anthropic has since shipped intervening 4.6-4.8 generations and Opus 5 (2026-07-24), with a published migration path off 4.5; no retirement date has been announced. Both Alibaba's Qwen3-VL (announced 2025-09-22, core weights released through 2025-10-21) and Zhipu/Z.ai's GLM-4.5V (released 2025-08-11) are open-weight/API releases carrying no formal GA label at launch, and both already face documented supersession: Alibaba has scheduled essentially the entire Qwen3-VL Model Studio family -- flash, 8B, 30B-A3B, 32B and 235B-A22B, instruct and thinking variants alike -- for retirement on 2026-10-10 in favor of Qwen3.6/3.7, while Artificial Analysis already labels GLM-4.5V "deprecated" in favor of Z.ai's own GLM-4.6V (released 2025-12-08) and GLM-5V-Turbo (2026-04-01/02), even though Z.ai's live API documentation still serves the 4.5V endpoint with no retirement notice of its own. On independent human-preference ranking (Arena Vision Leaderboard, 2026-08-01), the four separate sharply: Gemini 3 Pro Preview ranks 11th (1289, an archived listing given the shutdown), Qwen3-VL-235B-Instruct ranks 63rd (1215) with its Thinking sibling trailing at 77th (1189), and GLM-4.5V trails the group at 96th (1156); Claude Opus 4.5 was not present in the cited Vision Arena snapshot. None of the four rankings reproduces its own vendor's launch-table leadership claim. Independent hallucination measurement tells a different, cautionary story: Artificial Analysis's AA-Omniscience found Gemini 3 Pro at 88% and Claude Opus 4.5 at 58%, both far higher than either vendor's headline capability numbers would suggest. On a foundational visual-reasoning benchmark designed to test child-level perception rather than expert task completion (BabyVision, revised 2026-07-07, external preprint), the three tested models scored far below the study's 94.1% adult baseline -- Gemini 3 Pro Preview highest at 49.7%, Qwen3-VL variants at 19.2-22.2%, and Claude Opus 4.5 lowest at 14.2%; GLM-4.5V was not included in that study. Where GLM-4.5V and Qwen3-VL were tested on free-form geometric/architectural-drawing grounding (CrossProjection, 2026-08-01), both scored reasonably on closed-choice categorical questions (57-62%) but collapsed on candidate-free point and line localization (0-36% PCK), a gap the authors describe as an initial diagnostic rather than a final verdict. Agentic computer-use evidence is thinnest and most partner-mediated: Anthropic's own computer-use tool remains in beta with an explicit prompt-injection warning, and the only OSWorld-leading claim for Opus 4.5 comes from partner UiPath's own scaffold rather than Anthropic's bare-model score; an independent test of Qwen3-VL-30B on screenshot-only OSWorld tasks (Hanyang University, 2026-07-30) found only about 28% success with little benefit from longer step budgets. Named production evidence is real but thin and rarely independently audited across all four: Notion (Claude), Box / GitHub Copilot preview / Figma experimental with a February 2026 Google-side reliability incident (Gemini, several explicitly preview or alpha), Microsoft Azure AI Foundry distribution without named customers (Qwen3-VL), and Zhipu's AutoGLM 2.0 -- confirmed live on iOS/Android/web by a regulatory HKEX filing and independent Chinese media, but with no GLM-4.5V-specific usage volume disclosed.

Caveats

Independent Arena ranks and AA-Omniscience hallucination rates are live human-preference/evaluation snapshots (2026-08-01) over a changing competitor pool, not fixed task-ground-truth accuracy, and Claude Opus 4.5 was not present in the cited Vision Arena pull. Gemini 3 Pro's cited independent scores come from a pre-release Artificial Analysis build of a production preview endpoint that Google has since shut down (2026-03-09); treat any "currently deployable" framing for it as archived, not live. BabyVision's child-comparison sample is drawn from one school population, and its Qwen3-VL row carries an Alibaba-affiliated co-author plus a Qwen3-Max judge model, so it is mixed-affiliation rather than fully independent for that model specifically. CrossProjection and the Hanyang University OSWorld study are each single, recent (2026-07/08) preprints, not yet peer-reviewed. UiPath's OSWorld-Verified claim for Claude Opus 4.5 is a partner-run scaffold result, not Anthropic's own bare-model reproduction. Vendor benchmark-leadership claims for all four models (Anthropic's OSWorld/MMMU figures, Google's launch table, Qwen's "led most metrics" claim, and Z.ai's 42-benchmark SOTA claim) remain vendor-run and have not been independently reproduced under identical settings. Re-date this cluster on the scheduled Qwen3-VL Model Studio retirement (2026-10-10) or any further Claude/GLM lifecycle change.

Approved by pragma.vision editorial — standing authorization (operator dev@soft.house, 2026-07-17) · 2026-08-05

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.