Verification register AI & Agents
Current reading
Readiness band and full integer triple
AI-assisted assembly · derived results
Readiness band
Mature
Primary summary from verified readiness
- Confidence
- 75% · stale
- Computed at
- 2026-08-20T10:54:06.022525+00:00
- Claimed
- 90
- Reported
- 88
- Verified
- 83
- Gap
- +7
Public ambition and stated capability
Observed practitioner reporting
Independently supported evidence
Claimed leads verified
Evidence strength Strong
Decision
What the current evidence supports
Human editorial judgment · 2026-08-20
Adopt with guardrails
- Why
- It is the de-facto primary academic benchmark for function calling, actively maintained (updated April 2026) with realistic agentic categories — reliable as a comparative signal, provided we account for the ~75% ceiling and the FC-vs-prompt methodology split rather than treating scores as production readiness.
- Next
- Use BFCL V4 as one input when choosing a tool-calling model for agent work, but pin the commit/package version and read sub-category scores (multi-turn, agentic, format sensitivity) rather than just the headline number; pair with our own task-representative evals.
Constraints
Blockers
No named blocker is present in the current public projection.
Evidence summary
Derived counts
AI-assisted assembly
- Total
- 12
- Tier 1
- 0
- Tier 2
- 4
- Tier 3
- 8
- Supports
- 4
- Contradicts
- 2
- Context
- 6
- Latest observed
- 2026-06-01
Counts and dates only. Raw signals, private excerpts, trust records, and internal corpus material are not published here.
Publication record
Revisions
Initial public reading
- 2026-07-19 Reading moved from mature to mature.