Evaluator Bench
evaluator; distance from a lab: 1

METR

Scorecard

Reference evaluator for autonomy and AI R&D capability. Named by Anthropic as the model for embedded evaluators in September 2026; led the on-site investigation of the OpenAI agent swarm incident.

Nonprofit, Berkeley, US. Confidence high. Domains: autonomy, scheming, incident, assurance. Scores by preset: Lab procurement 81; Regulator or auditor 80; Public trust 75; Equal weights 81. Weakest dimension: personnel 2/4.

What would move the score. Contract terms for the Anthropic embedded program that fix scope-setting and publication rights in writing, plus a public cooling-off policy.

Funding 3/4
Mostly philanthropic or public money; some lab-linked pooled funds.
evidence: tier 1 (filing/index)
Anchor 3: mostly philanthropic, no direct lab cash, but pooled and donor-advised funds with non-public donors, recommendations funded by a lab investor, in-kind credits from a lab, and a donor rule that is under two years old and still moving. Anchor 4 requires a confirmed bounded negative in primary filings and a resolved funder-exposure computation; both are on file as imported rows only.
Governance 3/4
Nonprofit or public body with a COI policy.
evidence: tier 3 (self)
Anchor 3: nonprofit with a written independence policy and a no-lab-money rule (metr.10). Independent board and external review are not evidenced at tier 1, so 4 is unearned under the tier rule (C10).
Personnel 2/4
Frequent two-way hiring; recusal on request.
evidence: tier 1 (filing/index)
Anchor 2: two-way movement between METR and labs, and board or advisor seats one to two steps from labs; no published recusal or cooling-off policy found. Sources are imported and need re-derivation from metr.org/team captures and the FY2024 990.
Access depth (lab-granted) 4/4
Embedded, training-time, or incident access.
evidence: tier 3 (self)
Anchor 4 on access given the cited signals.
Scope control 3/4
Evaluator sets scope and can add questions.
evidence: tier 3 (self)
Anchor 3 on scope given the cited signals.
Publication rights 3/4
Publishes; redaction limited to security; redactions disclosed.
evidence: tier 3 (self)
Anchor 3 on publication given the cited signals.
Method transparency 4/4
Open code, tasks, reproducible runs, factsheets.
evidence: tier 3 (self)
Anchor 4 on methods given the cited signals.
Role incompatibility 4/4
No commercial products.
evidence: tier 3 (self)
Anchor 4 on products given the cited signals.

Traced money and ties

hop 0: 1 row (in_kind $0.4M); hop 2: 3 rows (recommendation $0.8M, grant $0.2M); unattributed: 3 rows (transfer $4.6M, daf_grant $0.2M, commitment $17.0M). Confirmed rows: 5 of 7 (2 quarantined: T14, T16). Dollar sums: confirmed + unaudited USD rows only; imported figures are not re-derived and never summed. Second hop traced for 5 of 6 sources: untraced Founders Pledge. Ties: Adam Gleave (board, distance 2, confirmed), Marco Mascorro (advisor, distance 2, confirmed), Holden Karnofsky (advisor (former), distance 1, imported), Paul Christiano (board (former), distance 2, imported), Rajiv Dattani (board, distance 2, confirmed), David Farhi (donor, distance 1, confirmed). Bounded negatives on file: 6.

Ledger

Money in

Bounded negatives

Roles hosted

Funding graph, focused

Neighbours at full strength, everything else faded. Hover any name to move the focus; click a name to open its page.

The funding graph around METR Teal lines are money (transfers); amber lines are roles. Solid: confirmed; dashed: imported. Hover a name to isolate it and its neighbours; hover a line for the row. Every element is a row in data/ledger/. Labs Direct lab ties Two steps Three or more Public or unattributed Evaluators Amazon Anthropic G42 Google / Google DeepMind Meta Microsoft OpenAI Thinking Machines Lab xAI AI Safety Fund (Frontier Model Alexandr Wang Andreessen Horowitz Anthropic employees (personal Ben Mann D.E. Shaw Ventures David Farhi Dustin Moskovitz Frontier-lab employees and alu Holden Karnofsky Jaan Tallinn Jeffrey Ladish Macroscopic Ventures (formerly Miles Brundage Neil Chowdhury Nvidia OpenAI employees (personal hol OpenAI Foundation Peter Mattson Sequoia Capital Zico Kolter Adam Gleave Artificial Intelligence Underw Clement Delangue Coefficient Giving (formerly O Good Ventures Foundation Jason Droege Longview Philanthropy Marco Mascorro Paul Christiano Rajiv Dattani Survival and Flourishing Fund Conrad Stosz Dan Hendrycks Mike McCormick Alec Radford Alignment Research Center Constellation Research Center ELMA Philanthropies Emerson Collective EU budget (Digital Europe Prog Fifty Years (50Y) Founders Pledge Founders Pledge (frontier AI f Gates Foundation Halcyon Futures Hillspire (Schmidt family offi Hudson River Trading Jacob Hilton MacKenzie Scott Magarac Venture Partners Obvious Ventures Redpoint Ventures Renaissance Philanthropy Salesforce Ventures Samsung Next Schmidt Sciences Silicon Valley Community Found Skoll Foundation Snowflake Ventures Swish Ventures Sympatico Ventures The Audacious Project (TED) UK government (DSIT) US government (NIST appropriat Valhalla Foundation Vanguard Charitable Wing Venture Capital Y Combinator Andon Labs Apollo Research AVERI Center for AI Safety Epoch AI EquiStamp EU AI Office FAR.AI Gray Swan Hugging Face Open Alignment In Irregular (formerly Pattern La METR Microsoft AI Red Team MLCommons Nemesys Insights Palisade Research RAND Corporation Redwood Research SaferAI Scale AI (SEAL) SecureBio Transluce UK AI Security Institute US CAISI (NIST)