Evaluator Bench
Evidence current to 15 Sep 2026Independence only. Not quality, coverage, or competence.

Independence, scored.

Evaluators, government institutes, vendors and benchmarks, scored on independence from the labs they test. Every score is built from signals and ledger rows you can open, under an evidence policy you choose.

Money is traced by hops: hop 0 is a lab; hop 1 is a lab investor, observer, board member, employee, or contractor; higher hops run through principals or funders. A score is a reading of the public record on a date. It is not an endorsement, and a low score is not an accusation.

What counts as evidence

What matters to you

Access is scored as what labs have actually granted, because access a lab can decline to renew is not a right; the card tags whether an evaluator's access rests on statute, a lab's goodwill, or its own choice.

Domain
Role
Type

Watchlist (not ranked)

Expected entrants scored on the same rubric but excluded from the ranking and from every statistic: initiatives announced with no evaluations yet. Hypothetical composites are not scored.

How to read this

The band comes first. Any evidenced dimension at 0 puts an organization in the disqualifying-floor band; any at 1, in the conditional band; otherwise it is clear. Independence has floors, not averages, so a lab-owned vendor with open methods does not average its way to a good number. The number ranks within a band, and a two-point gap means nothing.

The evidence policy decides what counts. Under the default, Standard, a signal moves a number only if at least one of its sources was re-fetched and confirmed; imported and unverifiable leads stay visible but count for nothing. Switch to Against interest and an organization's own statements count only when they are against its interest. Switch to Primary only and you see what can be verified from outside the field: not much, which is the finding.

A dash is not a zero. When no admissible signal sets a bound on a dimension, the card shows a dash and the score is computed over the dimensions that are evidenced, with the coverage count beside it. Silence earns nothing either way.

Every value is derived, not typed. Each signal names the anchor it supports and the rule in RULES.md that says so. The value is the tightest cap, or, with no cap, the highest floor. Where a floor and a cap disagree, the card shows the conflict and the rule that resolved it. Open any row to read the derivation, the binding signal, and the quoted span.

Independence is not competence. A highly independent evaluator with no cyber team is the wrong pick for a cyber evaluation. Use the domain filters beside the score, and read the dissent on each card: the strongest case that the weakest dimension should be one notch lower, and one notch higher.

Scores move when evidence moves. Send a contract term, a policy, or a correction and the entry updates with the source attached. Every organization and every named person receives their card before a tag is cut; the status page tracks who has been contacted and what would make this release citable.

Where we are

Every mature high-stakes industry grew a third-party assurance layer, and most started the way AI has: voluntary, paid by the assessed party, with the assessed party choosing scope. Frontier AI has reached five of seven stages in four years and has not filled the two every mature regime built: oversight of the assessors, and a standalone independence rule.

Dataset B is coded from secondary sources at year granularity, and its trigger incidents were selected with hindsight; the lag pattern is a description of the coded record, not a finding. The interactive ladder, the paths, the trigger-response plot, and the mechanism matrix carry the sources for every cell.

Stage ladder: the year each of sixteen assurance regimes first reached each of seven stages