pragma.vision Technology observatory

Verification register Security & Identity

Living definition

What is METR RE-Bench?

AI Safety, Eval & Alignment Security & Identity

As of
2026-07-23
Revision
2026-07-26.0
Method
v1.3.0

Definition

The term, in context

AI-assisted draft · approved dataset

RE-Bench is an evaluation suite from METR that measures how well autonomous agents perform open-ended machine learning research engineering work, placing them in environments with an explicit scoring function and a fixed time budget in which to improve a solution. Human experts attempt the same environments under comparable conditions, so agent scores can be read against a human baseline rather than an absolute threshold.

Live readiness status

Status as of the current dateline

AI-assisted assembly · derived results

As of 2026-07-23, verified readiness is 70 (claimed 80, reported 75, gap 10) — Strong evidence strength; current signals suggest Adopt with guardrails.

The readiness fact belongs to the canonical Readiness Verdict for METR RE-Bench.

Dataset approval

Human editorial release

pragma.vision editorial — standing authorization (operator dev@soft.house, 2026-07-17)

Dataset
2026-07-26.0
Hash
sha256:e2eb667df093db8819d579dff9aad51021b4977da85db2dfa82f63f9ed4e72b7
Approved
2026-07-26

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.