Independence, scored.
Evaluators, government institutes, vendors and benchmarks — plus a watchlist of expected entrants — scored on independence from the labs they test. Every score is built from public signals and ledger rows you can inspect.
Money is traced by hops: hop 0 is a lab; hop 1 is a lab investor, observer, board member, employee, or contractor; higher hops run through principals or funders. "Lab-tied" means hop 0 or hop 1. Scores are gated ordinal projections, not procurement truth.
Set the weights for your situation; the ranking updates as you move them.
What matters to you
Watchlist (not ranked)
Expected entrants scored on the same rubric but excluded from the ranking and from every average: a hypothetical composite and an initiative announced with no evaluations yet.
The rubric
Eight dimensions, each scored 0 to 4 from public evidence, then weighted. The dimensions follow the AI Evaluator Forum's AEF-1 operating conditions and the financial-audit independence rules that Illinois SB 315 imports for frontier AI, with two additions the field tends to skip: who owns the evaluator, and whether it sells fixes to the companies it grades.
Signals
What raises and lowers an independence score. The list is deliberately concrete: each item is something you can verify from a filing, a contract term, a system card, or a published policy.
Raises the score
- A written policy refusing money from frontier developers, including donations directed by their staff.
- A published conflict-of-interest policy: no outcome-contingent fees, no side grants or investments from labs being evaluated, mandatory recusal for anyone with a financial interest.
- No single funder above a stated share of budget; funders disclosed by name.
- Publication rights fixed in the contract before work starts, with redaction limited to security-sensitive or privileged material and a public statement of what was redacted.
- A record of publishing findings the lab did not like: scheming, reward hacking, shutdown resistance, safeguard failures.
- Open evaluation code and task suites so results can be re-run by others.
- On-site, weights-level, or training-time access rather than a few weeks on an API.
- Access that does not depend on the lab's goodwill: statute, regulation, or a court-enforceable agreement.
- Membership in a standards body (AEF-1, AVERI pilots) and disclosure of operating conditions for each evaluation.
- Cooling-off periods before staff move to labs; equity in labs disclosed or prohibited.
- A client base spread across many developers, so no lab is a dominant revenue source.
- Double-blind or secure-enclave protocols that let the evaluator work without the lab seeing the test items.
Lowers the score
- The lab pays for the evaluation of its own model. Common, and the single most under-discussed conflict in the field.
- Investors shared with the labs: a venture firm that backs both the evaluator and OpenAI, or a chip vendor buying the evaluator while anchoring a lab's IPO.
- Majority ownership by a frontier developer, or status as a business unit of one.
- Leadership holding a board seat, safety-committee chair, or advisory role at a lab the organization evaluates, even with recusal.
- Selling defenses, guardrails, or monitoring products to the same labs it evaluates: the audit-plus-consulting problem that Sarbanes-Oxley separated in accounting.
- A contract that forbids disclosing who funded a benchmark or evaluation.
- The lab sets the scope, the time window, and which questions are out of bounds.
- The lab reviews drafts for tone and emphasis, not only for security redactions.
- Only aggregated or lab-summarized results reach the public.
- Access that the lab can decline to renew with no consequence.
- Political or budgetary dependence that can redirect a government institute's mandate within a year.
- Co-authoring research with a lab while also serving as its external evaluator.
- Free tokens and compute from the lab, when they are a material share of operating capacity.
- Heavy talent flow in both directions between the evaluator and the labs.
- No track record: a pledge to evaluate is not an evaluation.
How good is the evidence
The scores measure what the public record shows, and the public record is largely what the evaluators say about themselves. Among the 112 unique sources cited by 295 signals for the ranked population, 60 (54%) are tier-3 self-published and 10 (9%) are tier-1 sources (regulatory filings or public indexes). 41 of 295 signals (14%) carry an exact quoted span. The evidence mix varies: AVERI has 5 self-published sources out of 7, with 0 tier-1; SaferAI 2 of 5, with 1 tier-1; and METR 7 of 21, with 5 tier-1. An evidence-based score rewards silence, and self-published sources can be replaced only where independent reporting exists — which, for most evaluators, it does not. That scarcity is a finding, not a data gap to be patched: the field's independence cannot be verified from outside the field. Scores built mostly on self-report are flagged by their evidence tier on every scorecard.
Cases that set the bar
Six engagements from the last two years, read for what they reveal about each dimension. Longer treatments and sources are in the repo under paper/.
Where we are
Every mature high-stakes industry grew a third-party assurance layer, and most of them started the way AI has: voluntary, paid by the assessed party, with the assessed party choosing scope. The ladder marks the year each regime first reached each stage. Tap a cell to see the milestone behind it and its source; tap a regime name for its full history.
A caveat the ladder cannot show: it records first arrivals, so a one-clause rule in one state fills a cell the same way Sarbanes-Oxley does, and it hides the order things happened in. The three views below fix both.
Read financial audit left to right and you get 158 years from the first mandate to the first rule separating auditors from consultants. Read the AI row and you get four. Across regimes, rules have followed public failures within a few years in the coded record (triggers were selected with hindsight, so this is a description, not a causal finding); the stretch before the first failure is where a regime's shape is decided. The oversight and independence-rule columns are the ones AI has not filled. Data and method: data/industries/, bench/timeline.py, paper/.
Paths, responses, and mechanisms
Three views that keep the nuance the ladder drops. Hover or tap any glyph, point, or cell to read what it stands for and where it comes from.
The order, not the calendar
Each regime as a sequence of moves. Reversals show up in red, the assessed party absorbing the assessment function in amber. The right column ranks regimes by how closely their opening moves match frontier AI's so far, using edit distance over the sequences.
Fast is not the same as strong
Every trigger incident paired with the next rule that followed. Most responses arrive within three years, and most are mandates or standards rather than structural changes; the two structural responses in the set are the PCAOB after Enron and lab inspection under Good Laboratory Practice after the IBT fraud. Frontier AI's two responses so far score 2 and 1.
Who pays, selects, sees, publishes, and watches
The configuration the ladder cannot show. Read the AI row against the others: the only other regime with the same five-cell pattern is dietary supplements, the regime that removed pre-market review by statute. The match is on a coarse categorical coding, so it is suggestive rather than conclusive.
Money and ties, every evaluator
The ledger view: each evaluator's traced inflows by distance from a frontier lab, ties within two steps, and how much of the record has been re-derived from primary sources rather than imported. Read the confirmed column before the others; most rows are still leads. Rows and sources: data/ledger/. Rebuild with python -m bench exposure --json.
Two patterns worth reading off the matrix. First, the two organizations with the most complete self-disclosure, Transluce and AVERI, score lower on funding than several that disclose less; an evidence-based score rewards silence unless the empty cells are read as gaps, which is why the confirmed and second-hop columns sit beside the rows. Second, the same three or four funders (Coefficient Giving, the Survival and Flourishing Fund, the Audacious Project, the EU AI Office) sit behind most of the nonprofit evaluators, so a question about one evaluator's independence is often a question about one funder's.
The graph itself
Every entity in the ledger placed by its distance from a lab, with money in teal and roles in amber. Dashed lines are imported rows not yet re-derived. Hover a name to isolate its neighbourhood; click it to open the entity page and walk the graph hop by hop.
Browse by type
Every item has its own page. Dockets: contestable-claim drafts for future Epistemedia review (none submitted; drafts are not evidence). Status: coverage, confirmed shares, open questions, data downloads. Evaluators: the scorecard table and a page per organization with its dimensions, signals, sources, ledger rows and a focused funding graph. Entities: every funder, lab, investor and person in the ledger, with money in and out and roles held. Regimes: the sixteen assurance regimes with milestones, mechanisms and path signatures. Sources: each cited source with its tier, audit status, and everything that cites it.
How the scores are made
Each evaluator gets a 0 to 4 on eight dimensions using only public evidence: filings, funding announcements, system cards, published policies, contracts described in reports, and press. The weighted total is scaled to 100. Scores are gated by default: any 0 on a dimension caps the total at 40, any 1 caps it at 60 — independence has floors, not just averages. Uncheck the gate for the raw compensatory score (sensitivity analysis). The lab-procurement preset is the pre-registered confirmatory view; the other presets are sensitivity checks, not alternative truths. Confidence tags flag where evidence is thin. An entry is not an endorsement, and a low score is not an accusation; it means the public record does not yet show the safeguards that would earn a higher one.
Independence is one axis. Competence, domain coverage, staffing, and turnaround are others, and a highly independent evaluator with no cyber team is the wrong pick for a cyber evaluation. Use the domain filters alongside the score.
Scores move when evidence moves. Send a contract term, a policy, or a correction and the entry updates with the source attached. This dataset was built by a single curator with AI assistance (Claude, Anthropic — itself a frontier developer in the ecosystem scored here); curator disclosure, including the fields still awaiting confirmation, is in DISCLOSURE.md. Corrections received through 29 Sep 2026 are filed as signals with the date received, and evaluators whose scores changed after publication receive a right-of-reply packet before the next batch.
How to submit evidence Status and data Dockets