Living definition
What is METR RE-Bench?
AI Safety, Eval & Alignment Security & Identity
- As of
- 2026-08-20
- Revision
- 2026-08-17.1
- Method
- v1.3.0
Definition
The term, in context
AI-assisted draft · approved dataset
RE-Bench is an evaluation suite from METR that measures how well autonomous agents perform open-ended machine learning research engineering work, placing them in environments with an explicit scoring function and a fixed time budget in which to improve a solution. Human experts attempt the same environments under comparable conditions, so agent scores can be read against a human baseline rather than an absolute threshold.
Live readiness status
Status as of the current dateline
AI-assisted assembly · derived results
As of 2026-08-20, verified readiness is 70 (claimed 80, reported 75, gap 10) — Strong evidence strength; current signals suggest Adopt with guardrails.
The readiness fact belongs to the canonical Readiness Verdict for METR RE-Bench.
Dataset approval
Human editorial release
pragma.vision editorial — standing authorization (operator dev@soft.house, 2026-07-17)
- Dataset
- 2026-08-17.1
- Hash
- sha256:0c37108a182f510fc8cf4d1dfc3a30339f2b59e62654201edea2f530a268707e
- Approved
- 2026-08-17