How the scores are made
Values
Each evaluator gets a 0 to 4 on eight dimensions using only public evidence: filings, funding announcements, system cards, published policies, contracts described in reports, and press. Every signal names the anchor it supports and the rule in RULES.md that says so. The value on a dimension is the tightest admissible cap or, with no cap, the highest admissible floor. Where a floor and a cap disagree, the assessment carries a written resolution naming the rule, and the card shows the conflict. A 0 needs a quoted span; a 4 needs a span and a tier-1 or tier-2 source or two independent sources with one not self-published; a bound from sources that were not all confirmed cannot set an extreme. Bounds that fail these tests are held at the nearest supportable anchor and the card says which rule held them.
Evidence policies
The reader chooses what counts. The default, standard, admits a signal only if at least one cited source was re-fetched and confirmed, so no number moves on a source you cannot open and check. Imported and unverifiable leads stay visible and count for nothing. A dimension with no admissible signal renders as a dash and is excluded from the score; the coverage count sits beside every score.
| Evidence policy | Rule | Ranked signals that count | Assessments unevidenced |
|---|---|---|---|
| Leads included | Every signal that is not quarantined counts, including imported and unverifiable leads, which can support only a 2. For research, not for citation. | 317 of 318 | 1 of 192 |
| Standard (default) | A signal counts only if at least one cited source was re-fetched and confirmed. The default: no number moves on a source you cannot open and check. | 303 of 318 | 6 of 192 |
| Against interest | Standard, and an organization's own statements count only when they are against its interest or backed by a non-self source. | 225 of 318 | 48 of 192 |
| Verified spans | Standard, and the signal carries a quoted span you can find on the page. | 267 of 318 | 25 of 192 |
| Primary only | Only confirmed filings, funder indexes, or third-party ledgers with a quoted span: what can be verified from outside the field. | 26 of 318 | 171 of 192 |
Bands and scores
The weighted total over the evidenced dimensions is scaled to 100 under four weight presets, each with a written derivation. Independence has floors, so the band comes first: an evidenced 0 on a conflict dimension (funding, governance, personnel, role incompatibility, scope, publication) is a disqualifying floor, a 1 on one of those a conditional floor, otherwise clear. Access and methods count in the number and never set a band, because a 0 there means the labs have not let the organization in or it has not published its methods, not that it is compromised. The number ranks within a band, and on the directory it is hidden until the reader chooses weights, so two readers with different priorities see different numbers and neither is the site's. The lab-procurement preset is the confirmatory view; it was fixed in the seed script before the population pass but not registered outside this repository. The other presets are sensitivity checks, not alternative truths. Scores are ordinal projections over anchored rubrics: a ten-point gap is not twice the independence, and a two-point gap is nothing.
Independence is one axis
Competence, domain coverage, staffing, and turnaround are others, and a highly independent evaluator with no cyber team is the wrong pick for a cyber evaluation. Use the domain filters alongside the score. Government bodies are scored on what reaches the public and what access they hold, with the mechanism tagged statutory where the constraint is the law rather than a lab. An entry is not an endorsement, and a low score is not an accusation; it means the public record does not yet show the safeguards that would earn a higher one.
How good is the evidence
The scores measure what the public record shows, and the public record is largely what the evaluators say about themselves. Among the 181 unique sources cited by 318 signals for the ranked population, 100 (55%) are tier-3 self-published and 21 (12%) are tier-1 sources (regulatory filings or public indexes). 275 of 318 signals (86%) carry an exact quoted span. Under the standard evidence policy 303 of 318 signals count; under the primary-only policy 26 do. That scarcity is a finding, not a data gap to be patched: the field's independence cannot yet be verified from outside the field. An evidence-based score would reward silence if silence defaulted to a number; here an unevidenced dimension renders as a dash and is excluded from the score, and every card shows the evidence tier of what it rests on.
Who made this, and what they hold
Curated by Yohei Nakajima: managing partner at Untapped Capital, a pre-seed and seed venture fund; operator of Epistemedia, the claim-adjudication layer this repository drafts dockets into; author of ActiveGraph, the runtime the build runs on. The site is a collaboration between the curator and models from several developers: the first draft of the curation and the rules pass were written by Claude (Anthropic); Codex (OpenAI) and Grok (xAI) ran independent verification passes recorded in paper/audits/; Gemini (Google) and Muse reviewed and criticized the site and the paper. Anthropic, OpenAI, xAI, and Google are labs in this ledger, evaluated by organizations scored here. On the concern that Anthropic's model produced a ranking with METR, the evaluator Anthropic named in its September 2026 commitment, at the top: every value is derived from public bounds under rules applied to every organization the same way, the ranking is re-derivable by anyone with any tool or none, and the against-interest and primary-only views are one click away; if a model's involvement had tilted a value, it would show as a bound or a rule, both open to correction. Holdings: small public-market shares in Google and Meta, and a private holding in SpaceX, which owns xAI; assessments of evaluators with confirmed ledger ties to those labs say so in their rationale. Shared funders between Untapped Capital and the evaluators' funders: not yet checked; the check is scheduled before the freeze and its result will replace this sentence. One coder; a second coder on every extreme is a gate for a citable tag. The full statement is in DISCLOSURE.md.
Scores move when evidence moves
Send a contract term, a policy, or a correction and the entry updates with the source attached. Every ranked organization and every named person receives their card and a fourteen-day reply window before a tag is cut; the status page tracks who has been contacted. After first publication, a score that moves by more than one anchor triggers a fresh record packet before the next tag.
How to submit evidence