Evaluators
26 ranked organizations scored on eight independence dimensions, split by role (referee, government, vendor, benchmark, lab team). Columns show the weighted score under each preset, the weakest dimension, confidence, and how many ledger inflow rows are confirmed. Open a row for the full scorecard, ledger, and focused graph.
| Evaluator | Role | Lab | Regulator | Public | Equal | Floor | Confidence | Confirmed rows | Domains |
|---|---|---|---|---|---|---|---|---|---|
| METR Nonprofit, Berkeley, US | referee | 81 | 80 | 75 | 81 | personnel 2 | high | 5/7 | autonomy, scheming, incident, assurance |
| AVERI Nonprofit, Washington, US | referee | 79 | 78 | 78 | 81 | funding 3 | med | 4/4 | assurance |
| SaferAI Nonprofit, Paris, FR | referee | 76 | 79 | 80 | 75 | access depth (lab-granted) 2 | med | 2/3 | assurance, cyber, bio, misuse |
| EU AI Office Government, Brussels, BE | government | 75 | 74 | 73 | 75 | publication rights 2 | med | 1/1 | bio, cyber, misuse, assurance |
| FAR.AI Nonprofit, Berkeley, US | referee | 75 | 75 | 75 | 75 | funding 3 | high | 4/6 | jailbreak, bio, cyber, misuse |
| UK AI Security Institute Government, London, UK | government | 75 | 74 | 73 | 78 | publication rights 2 | high | 2/3 | cyber, bio, jailbreak, autonomy, assurance |
| Redwood Research Nonprofit, Berkeley, US | referee | 74 | 73 | 73 | 75 | personnel 2 | med | 1/2 | scheming, incident, assurance |
| Transluce Nonprofit, Berkeley, US | referee | 73 | 75 | 73 | 75 | funding 2 | high | 3/3 | misuse, autonomy, benchmarks, assurance |
| RAND Nonprofit, Santa Monica, US | referee | 71 | 69 | 70 | 72 | publication rights 2 | med | 2/2 | bio, cyber, assurance |
| SecureBio Nonprofit, Cambridge, US | referee | 71 | 71 | 69 | 72 | funding 2 | high | 3/3 | bio, misuse |
| Apollo Research Public benefit corp, London, UK / San Francisco, US | referee | 67 | 69 | 64 | 66 | funding 2 | high | 2/2 | scheming, autonomy, assurance |
| Epoch AI Nonprofit, Remote | referee | 64 | 65 | 64 | 69 | funding 2 | high | 4/8 | benchmarks |
| Andon Labs VC-backed, Stockholm, SE / San Francisco, US | referee | 60 | 61 | 59 | 59 | funding 2 | low | 1/1 | autonomy, benchmarks |
| Center for AI Safety Nonprofit, San Francisco, US | referee | 60 | 60 | 60 | 60 | personnel 1 | med | 2/2 | benchmarks, misuse |
| US CAISI (NIST) Government, Gaithersburg, US | government | 60 | 60 | 60 | 60 | publication rights 1 | high | 1/2 | cyber, bio, misuse, assurance |
| Holistic Agent Leaderboard (Princeton) Academic, Princeton, US | benchmark | 60 | 60 | 60 | 60 | access depth (lab-granted) 1 | med | 0/0 | benchmarks, autonomy |
| Humane Intelligence Nonprofit, New York, US | referee | 60 | 60 | 60 | 60 | access depth (lab-granted) 1 | med | 0/0 | misuse, jailbreak |
| Palisade Research Nonprofit, Berkeley, US | referee | 60 | 60 | 60 | 60 | access depth (lab-granted) 1 | med | 1/2 | cyber, autonomy, scheming |
| EquiStamp Private company, capital structure undisclosed, Remote | referee | 59 | 59 | 59 | 56 | governance 2 | low | 1/1 | assurance, autonomy |
| MLCommons (AILuminate) Consortium, San Francisco, US | benchmark | 59 | 60 | 55 | 60 | funding 1 | med | 1/1 | benchmarks, assurance |
| Irregular (formerly Pattern Labs) VC-backed, Tel Aviv, IL / San Francisco, US | vendor | 47 | 49 | 40 | 44 | funding 1 | high | 2/2 | cyber, misuse |
| Nemesys Insights Private company, capital structure undisclosed, Albany, US | vendor | 46 | 46 | 44 | 47 | funding 1 | low | 2/2 | bio, misuse, assurance |
| Dreadnode VC-backed, Remote | vendor | 42 | 43 | 34 | 38 | funding 1 | low | 0/0 | cyber |
| Microsoft AI Red Team Big Tech unit, Redmond, US | lab-team | 41 | 43 | 31 | 38 | funding 1 | med | 1/1 | jailbreak, cyber, misuse |
| Gray Swan VC-backed, Pittsburgh, US | vendor | 40 | 40 | 35 | 38 | role incompatibility 0 | high | 2/2 | jailbreak, cyber, misuse |
| Scale AI (SEAL / Scale Labs) VC-backed, San Francisco, US | vendor | 38 | 40 | 26 | 34 | funding 0 | high | 1/1 | benchmarks, cyber, bio, misuse |
Watchlist (not ranked)
Expected entrants scored on the same rubric but excluded from rankings and averages: hypothetical composites and announced initiatives with no evaluations yet.
| Evaluator | Role | Lab | Regulator | Public | Equal | Floor | Confidence | Confirmed rows | Domains |
|---|---|---|---|---|---|---|---|---|---|
| Assurance firms (Big Four and peers) Hypothetical composite, not a real organization, Global | expected-entrant | 38 | 39 | 41 | 38 | access depth (lab-granted) 1 | low | 0/0 | assurance |
| Hugging Face Open Alignment Initiative VC-backed, New York, US / Paris, FR | expected-entrant | 37 | 38 | 40 | 40 | governance 0 | low | 1/1 | assurance |