Evaluator Bench

Evaluators

26 ranked organizations scored on eight independence dimensions, split by role (referee, government, vendor, benchmark, lab team). Columns show the weighted score under each preset, the weakest dimension, confidence, and how many ledger inflow rows are confirmed. Open a row for the full scorecard, ledger, and focused graph.

EvaluatorRoleLabRegulatorPublicEqualFloorConfidenceConfirmed rowsDomains
METR
Nonprofit, Berkeley, US
referee81807581personnel 2high5/7autonomy, scheming, incident, assurance
AVERI
Nonprofit, Washington, US
referee79787881funding 3med4/4assurance
SaferAI
Nonprofit, Paris, FR
referee76798075access depth (lab-granted) 2med2/3assurance, cyber, bio, misuse
EU AI Office
Government, Brussels, BE
government75747375publication rights 2med1/1bio, cyber, misuse, assurance
FAR.AI
Nonprofit, Berkeley, US
referee75757575funding 3high4/6jailbreak, bio, cyber, misuse
UK AI Security Institute
Government, London, UK
government75747378publication rights 2high2/3cyber, bio, jailbreak, autonomy, assurance
Redwood Research
Nonprofit, Berkeley, US
referee74737375personnel 2med1/2scheming, incident, assurance
Transluce
Nonprofit, Berkeley, US
referee73757375funding 2high3/3misuse, autonomy, benchmarks, assurance
RAND
Nonprofit, Santa Monica, US
referee71697072publication rights 2med2/2bio, cyber, assurance
SecureBio
Nonprofit, Cambridge, US
referee71716972funding 2high3/3bio, misuse
Apollo Research
Public benefit corp, London, UK / San Francisco, US
referee67696466funding 2high2/2scheming, autonomy, assurance
Epoch AI
Nonprofit, Remote
referee64656469funding 2high4/8benchmarks
Andon Labs
VC-backed, Stockholm, SE / San Francisco, US
referee60615959funding 2low1/1autonomy, benchmarks
Center for AI Safety
Nonprofit, San Francisco, US
referee60606060personnel 1med2/2benchmarks, misuse
US CAISI (NIST)
Government, Gaithersburg, US
government60606060publication rights 1high1/2cyber, bio, misuse, assurance
Holistic Agent Leaderboard (Princeton)
Academic, Princeton, US
benchmark60606060access depth (lab-granted) 1med0/0benchmarks, autonomy
Humane Intelligence
Nonprofit, New York, US
referee60606060access depth (lab-granted) 1med0/0misuse, jailbreak
Palisade Research
Nonprofit, Berkeley, US
referee60606060access depth (lab-granted) 1med1/2cyber, autonomy, scheming
EquiStamp
Private company, capital structure undisclosed, Remote
referee59595956governance 2low1/1assurance, autonomy
MLCommons (AILuminate)
Consortium, San Francisco, US
benchmark59605560funding 1med1/1benchmarks, assurance
Irregular (formerly Pattern Labs)
VC-backed, Tel Aviv, IL / San Francisco, US
vendor47494044funding 1high2/2cyber, misuse
Nemesys Insights
Private company, capital structure undisclosed, Albany, US
vendor46464447funding 1low2/2bio, misuse, assurance
Dreadnode
VC-backed, Remote
vendor42433438funding 1low0/0cyber
Microsoft AI Red Team
Big Tech unit, Redmond, US
lab-team41433138funding 1med1/1jailbreak, cyber, misuse
Gray Swan
VC-backed, Pittsburgh, US
vendor40403538role incompatibility 0high2/2jailbreak, cyber, misuse
Scale AI (SEAL / Scale Labs)
VC-backed, San Francisco, US
vendor38402634funding 0high1/1benchmarks, cyber, bio, misuse

Watchlist (not ranked)

Expected entrants scored on the same rubric but excluded from rankings and averages: hypothetical composites and announced initiatives with no evaluations yet.

EvaluatorRoleLabRegulatorPublicEqualFloorConfidenceConfirmed rowsDomains
Assurance firms (Big Four and peers)
Hypothetical composite, not a real organization, Global
expected-entrant38394138access depth (lab-granted) 1low0/0assurance
Hugging Face Open Alignment Initiative
VC-backed, New York, US / Paris, FR
expected-entrant37384040governance 0low1/1assurance