pragma.vision Technology observatory & strategic foresight

Verification register AI & Agents

Readiness verdict

tau2-bench (τ²-Bench)

A dated reading of what is claimed, reported, and independently verified in the current evidence.

As of
2026-08-20
Revision
1
Method
v1.3.0

Current reading

Readiness band and full integer triple

AI-assisted assembly · derived results

Readiness band

Watch

Primary summary from verified readiness

Confidence
59% · stale
Computed at
2026-08-20T10:51:48.741002+00:00
Claimed
80

Public ambition and stated capability

Reported
68

Observed practitioner reporting

Verified
58

Independently supported evidence

Gap
+22

Claimed leads verified

Evidence strength Strong

Decision

What the current evidence supports

Human editorial judgment · 2026-08-20

Adopt with guardrails

Why
It is the de-facto industry standard for conversational tool-agent eval and its pass^k/dual-control framing maps onto our shared-world agent risks, but headline numbers are domain-, simulator- and harness-sensitive and must not be quoted bare.
Next
Use τ²-Bench's dual-control + pass^k methodology as the template for our customer-facing agent reliability gates; report pass^k (not pass^1), segment by domain difficulty, and pin a fixed user-simulator + harness version rather than citing a single leaderboard figure.

Constraints

Blockers

No named blocker is present in the current public projection.

Evidence summary

Derived counts

AI-assisted assembly

Total
10
Tier 1
0
Tier 2
6
Tier 3
4
Supports
2
Contradicts
7
Context
1
Latest observed
2026-06-01

Counts and dates only. Raw signals, private excerpts, trust records, and internal corpus material are not published here.

Publication record

Revisions

Initial public reading

  1. 2026-07-19 Reading moved from watch to watch.

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.