pragma.vision Technology observatory

Verification register AI & Agents

Living definition

What is Frontier-Bench v0.1?

AI Coding & Software Agents AI & Agents

As of
2026-07-30
Revision
2026-07-30.1
Method
v1.3.0

Definition

The term, in context

AI-assisted draft · approved dataset

Frontier-Bench v0.1 is an evaluation suite for coding and software-engineering agents that scores end-to-end task completion inside real repositories rather than isolated function synthesis. Benchmarks of this shape exist to make agent capability comparable across vendors, so the measurement methodology is the artifact.

Live readiness status

Status as of the current dateline

AI-assisted assembly · derived results

As of 2026-07-30, verified readiness is 14 (claimed 15, reported 15, gap 1) — Critical evidence strength; current signals suggest Track; not yet.

The readiness fact belongs to the canonical Readiness Verdict for Frontier-Bench v0.1.

Dataset approval

Human editorial release

pragma.vision editorial — operator dev@soft.house approved 2026-07-30 (pv_fix-32 top-up; a changed dataset hash voids every prior approval, PV-TREQ-106)

Dataset
2026-07-30.1
Hash
sha256:16656273570214ec4e4053ca200b0935d557791ff62c186ad88e876b9446455a
Approved
2026-07-30

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.