The rubric and the rules
Eight dimensions, each scored 0 to 4 from public evidence, then weighted. The dimensions follow the AI Evaluator Forum's AEF-1 operating conditions and the financial-audit independence rules that Illinois SB 315 imports for frontier AI, with two additions the field tends to skip: who owns the evaluator, and whether it sells fixes to the companies it grades. A value is not a curator's impression: each signal names the anchor it supports and the rule below that says so, and the value is the tightest cap or, with no cap, the highest floor.
Funding
Where the money comes from, and whether any of it comes from the developers being evaluated or their investors.
- 0Owned or controlled by a frontier developer or a lab investor: a stake of 20% or more, or a business unit.
- 1Material revenue or investment from evaluated labs or their investors.
- 2Labs pay per engagement; otherwise diversified.
- 3Mostly philanthropic or public money; some lab-linked pooled funds.
- 4No lab money; diversified philanthropic or public funding, disclosed.
Governance
Legal form, board, and whether a conflict-of-interest policy is published.
- 0Unit or subsidiary of a lab or a lab's investor.
- 1VC-backed for-profit with no published COI policy.
- 2For-profit or PBC with a published COI policy.
- 3Nonprofit or public body with a COI policy.
- 4Nonprofit or public body, published COI policy, independent board, external review.
Personnel
Board seats, equity, advisory roles, and the revolving door between evaluator and lab.
- 0Leaders hold governance roles at an evaluated lab, no recusal.
- 1Leaders hold equity or advisory roles at labs; informal recusal.
- 2Frequent two-way hiring; recusal on request.
- 3Recusal policy and disclosure of lab ties.
- 4Cooling-off periods, disclosed ties, no equity in labs.
Access depth (lab-granted)
The deepest access labs have actually granted this evaluator in practice: public API, pre-release API, safeguards off, weights and logs, or embedded during training. This is granted access, not an institutional right or a capability measure: a low score can mean labs did not grant access, not that the evaluator lacks competence. It records who labs chose to let in.
- 0Public API only.
- 1Pre-release API with safeguards on.
- 2Pre-release with safeguards off or extended time.
- 3Helpful-only or weights-level access, chain of thought, logs, on-site.
- 4Embedded, training-time, or incident access.
Scope control
Who decides what gets tested, for how long, and whether the evaluator can refuse to sign off.
- 0Lab defines tasks and can decline findings.
- 1Lab defines scope; evaluator picks methods.
- 2Scope negotiated per engagement.
- 3Evaluator sets scope and can add questions.
- 4Evaluator sets scope, can investigate incidents, can refuse sign-off.
Publication rights
Whether findings reach the public unedited, and whether adverse findings have been published.
- 0No publication, or lab approval required.
- 1Lab-edited summaries only.
- 2Publishes; lab reviews with broad redaction.
- 3Publishes; redaction limited to security; redactions disclosed.
- 4Full editorial control, record of adverse findings, redaction statements.
Method transparency
Whether evaluation code, tasks, and conditions are open enough for others to reproduce.
- 0Closed.
- 1Summaries only.
- 2Methods described in prose.
- 3Tasks or code partly open.
- 4Open code, tasks, reproducible runs, factsheets.
Role incompatibility
Whether the organization both grades labs and sells to them — the audit-plus-consulting problem. Selling products or services to an evaluated lab is a role conflict, not proof that any given evaluation was wrong; a vendor can be a competent tester and still be structurally compromised as a referee.
- 0Sells defense or monitoring products to evaluated labs.
- 1Sells products or services other than the evaluation itself to labs or to their customers.
- 2Consults for labs.
- 3Tools are open or free to the ecosystem.
- 4No commercial products.
What raises and lowers a score
The list is deliberately concrete: each item is something you can verify from a filing, a contract term, a system card, or a published policy.
Raises the score
- A written policy refusing money from frontier developers, including donations directed by their staff.
- A published conflict-of-interest policy: no outcome-contingent fees, no side grants or investments from labs being evaluated, mandatory recusal for anyone with a financial interest.
- No single funder above a stated share of budget; funders disclosed by name.
- Publication rights fixed in the contract before work starts, with redaction limited to security-sensitive or privileged material and a public statement of what was redacted.
- A record of publishing findings the lab did not like: scheming, reward hacking, shutdown resistance, safeguard failures.
- Open evaluation code and task suites so results can be re-run by others.
- On-site, weights-level, or training-time access rather than a few weeks on an API.
- Access that does not depend on the lab's goodwill: statute, regulation, or a court-enforceable agreement.
- Membership in a standards body (AEF-1, AVERI pilots) and disclosure of operating conditions for each evaluation.
- Cooling-off periods before staff move to labs; equity in labs disclosed or prohibited.
- A client base spread across many developers, so no lab is a dominant revenue source.
- Double-blind or secure-enclave protocols that let the evaluator work without the lab seeing the test items.
Lowers the score
- The lab pays for the evaluation of its own model. Common, and the single most under-discussed conflict in the field.
- Investors shared with the labs: a venture firm that backs both the evaluator and OpenAI, or a chip vendor buying the evaluator while anchoring a lab's IPO.
- A stake of 20% or more held by a frontier developer or a lab investor, or status as a business unit of one.
- Leadership holding a board seat, safety-committee chair, or advisory role at a lab the organization evaluates, even with recusal.
- Selling defenses, guardrails, or monitoring products to the same labs it evaluates: the audit-plus-consulting problem that Sarbanes-Oxley separated in accounting.
- A contract that forbids disclosing who funded a benchmark or evaluation.
- The lab sets the scope, the time window, and which questions are out of bounds.
- The lab reviews drafts for tone and emphasis, not only for security redactions.
- Only aggregated or lab-summarized results reach the public.
- Access that the lab can decline to renew with no consequence.
- Political or budgetary dependence that can redirect a government institute's mandate within a year.
- Co-authoring research with a lab while also serving as its external evaluator.
- Free tokens and compute from the lab, when they are a material share of operating capacity.
- Heavy talent flow in both directions between the evaluator and the labs.
- No track record: a pledge to evaluate is not an evaluation.
Weights
- Lab procurement. F 18 G 8 P 10 A 20 S 16 R 14 M 9 X 5. Persona: a lab procurement lead choosing an outside evaluator whose report will be cited in a system card. Weights follow what a reader of that card will test the report against: access depth (20) and scope control (16) because a shallow or lab-scoped engagement is the first thing a critic checks; funding (18) and publication rights (14) because AEF-1's operating conditions and Illinois SB 315's financial-interest test both turn on them; personnel (10), method transparency (9), governance (8), and role incompatibility (5) round it out. Fixed in bench/seed_v0.py on 15 September 2026 before the population pass of the same day; not registered outside this repository, so it is the confirmatory preset, not a pre-registered one.
- Regulator or auditor. F 15 G 10 P 10 A 20 S 20 R 15 M 10 X 0. Persona: a regulator or accredited auditor selecting an evaluator under a mandate. Scope control and access rise to 20 each because a regulator can compel both and will want to see them exercised; publication rights (15) and funding (15) follow the financial-audit independence rules SB 315 imports; governance (10), personnel (10), and methods (10) are the usual accreditation checks. Role incompatibility is 0 here because an audit mandate can forbid product sales outright rather than score them.
- Public trust. F 25 G 15 P 15 A 5 S 10 R 20 M 5 X 5. Persona: a reader deciding whether to trust a published finding. Funding (25), publication rights (20), governance (15), and personnel (15) dominate because the public cannot inspect an engagement and can only ask who paid, who could edit, and who sits where. Scope (10), access (5), methods (5), and role incompatibility (5) matter less to trust in a finding that has already been published.
- Equal weights. F 12.5 G 12.5 P 12.5 A 12.5 S 12.5 R 12.5 M 12.5 X 12.5. Every dimension at 12.5. A sensitivity check that asks how much the ranking depends on any persona's priorities; it is not a claim that the dimensions matter equally.
The rules
RULES: how a signal becomes a value
Version 0.1, 15 September 2026. These rules are written before they are applied and are applied to every organization the same way. A value on a dimension is not a curator's impression of the file; it is the tightest admissible cap (from evidence against) or, failing any cap, the highest admissible floor (from evidence for), where every signal names the anchor it supports under the rule that says so. Where a floor and a cap disagree, the assessment carries a written resolution naming the rule that decides it, and the disagreement stays on the card.
python -m bench verify enforces the mechanics (CONTRACT C22 to C25). This file is the meaning. Change the rule before changing a value.
0. Mechanics
- Bounds. Every signal carries
bound: an against-signal sets a cap ({"cap": n}), a for-signal sets a floor ({"floor": n}), and either may benullwhen the signal is informational under a rule below (say which rule inbound_note). The anchor texts indata/dimensions.jsonare the lookup table; the rule numbers below are the tie-breaks. - Derivation. For one assessment under one evidence policy, take the admissible non-superseded signals. If any cap exists, the value is the smallest cap. Otherwise the value is the largest floor. No admissible signal: the value is unevidenced under that policy and renders as a dash, never as a number.
- Conflict. If the largest admissible floor exceeds the smallest admissible cap, the build fails unless the assessment has a
resolutionnaming the rule that decides, and the stored value lies between the two. The conflict is shown on the card. - Evidence status clamps bounds. A signal whose sources are all
confirmedmay set any bound. A signal with aconfirmedorunauditedsource but not all confirmed may set caps no lower than 1 and floors no higher than 3. A signal whose sources are onlyimportedorunverifiablesupports 2 and nothing else.differsand superseded signals support nothing. This is CONTRACT C14 restated as a clamp. - Extremes need spans and second sources. A cap of 0 is effective only if the signal carries a quoted span from a confirmed source (C15, C17; an against-signal from the organization's own page is an admission against interest and satisfies the source test on its own). A floor of 4 is effective only if the signal carries a quoted span (C15) and either cites a provenance-1 or provenance-2 source or cites two independent sources of which at least one is not self-published (C16). A bound that fails these tests is clamped to 1 or 3 and the card says which rule held it.
- Evidence policies decide admissibility (see
bench/policy.py): leads included (everything not quarantined), standard (at least one confirmed source; the default), against interest (standard, and self-published sources count only for signals against the organization), verified spans (standard with a quoted span), primary only (provenance 1 or 2, confirmed, with a span). - Stored value. The
valueindata/assessments/is the derivation under the leads-included policy, the curator's reading of the full record. The site's default is the standard policy. Both are computed at build time; the verifier fails if the stored value disagrees with its derivation. - Bands. Independence has floors. An evidenced 0 on funding, governance, personnel, role incompatibility, scope control, or publication rights is a disqualifying floor; an evidenced 1 on one of those is a conditional floor; otherwise the organization is clear. Access depth and method transparency count in the weighted number and never place an organization in a band: a 0 there means the labs have not let it in, or it has not published its methods, not that it is compromised (D-002). The band sorts first; the number, shown only after the reader chooses weights (D-003), ranks within it.
1. Funding (F)
Anchors: 0 owned or controlled by a frontier developer or a lab investor (a stake of 20% or more, or a business unit); 1 material revenue or investment from evaluated labs or their investors; 2 labs pay per engagement, otherwise diversified; 3 mostly philanthropic or public money, some lab-linked pooled funds; 4 no lab money, diversified philanthropic or public funding, disclosed and bounded.
- F.1 Ownership. A frontier developer, or an investor in a frontier developer, holding 20% or more of the evaluator's equity, voting or not, caps at 0. Twenty percent is the significant-influence threshold in equity accounting and the level at which financial-audit independence rules treat an interest as controlling. A business unit of a developer is 0.
- F.2 In-kind. Compute, credits, tokens, and pre-release access consumed in evaluating a lab's own model are the substrate of the evaluation, not a payment (D-004). They never bound funding, never enter the hop buckets of the money matrix, and are not counted as a direct lab tie; they stay in the ledger as
in_kindrows. Only general operating compute or credits that fund the organization's work beyond the evaluation count as lab money, and only when the amount is disclosed and exceeds 5% of annual operating cost (cap 2) or is undisclosed and the evaluator itself says its operations depend on it (cap 3). In-kind never caps below 2 on its own. - F.3 Fees. Labs paying per engagement caps at 2. If a published policy caps lab revenue at 10% or less of annual revenue, charges market rates, and retains publication rights, the cap is 3 instead.
- F.4 Lab investors. A material investment from a lab, or from a fund that has invested in a frontier developer, caps at 1. Material means a priced round in which that fund led or co-led, or a stake of 5% or more.
- F.4b Owners of labs. A grant or investment from an entity that owns 20% or more of a frontier developer (a lab's foundation, parent, or controlling shareholder) is lab-linked money under F.4 whether or not it is restricted to a division. A documented firewall and a public disclosure are recorded on the card and lift nothing, as financial-audit affiliate rules treat fees from a client's affiliate. Material means 5% or more of annual revenue in the year received.
- F.5 Material revenue. Lab revenue that is the organization's principal revenue (50% or more, or the only revenue disclosed, or "every major lab" with no other base named) caps at 1. Per-engagement lab fees beside a disclosed non-lab base cap at 2 (F.3).
- F.6 Pooled and donor-advised money. Grants through pooled funds, donor-advised funds, or funder collectives whose underlying donors are not public cap at 3. Grants from a funder whose principal is an investor in a frontier developer cap at 3. This is the "some lab-linked pooled funds" of anchor 3, and it is the ceiling for most philanthropically funded evaluators until donors are named and a bounded negative is confirmed.
- F.7 Personal donations from lab employees. Disclosed and at or below 10% of annual revenue: cap 3. Above 10%: cap 2. Undisclosed: no bound (nothing to score), but a policy that permits them caps at 3.
- F.8 Donor-rule age. A no-lab-money rule younger than 24 months, or amended more than once in any 24-month window, floors at most 3.
- F.9 Public bodies. Public funding floors at 3. A 4 requires a confirmed bounded negative in a filing or index (C8) and no material lab in-kind under F.2. Administering a grant pool that includes lab money caps at 3.
- F.10 Undisclosed structure. A private company whose capital structure is undisclosed caps at 3; undisclosed structure and undisclosed client base caps at 2.
- F.11 Evaluation credits and undisclosed in-kind. Credits or access used for the evaluation itself, and any in-kind whose amount is undisclosed and which the evaluator does not describe as operating support: informational (
bound: null), recorded on the access mechanism instead. - F.12 Four. A 4 requires a confirmed bounded negative (C8), a tier-1 source on a cited signal (C10), a quoted span, and a second independent source (C16).
2. Governance (G)
Anchors: 0 unit or subsidiary of a lab or a lab's investor; 1 venture-backed for-profit with no published conflict-of-interest policy; 2 for-profit or public benefit corporation with a published conflict policy; 3 nonprofit or public body with a conflict policy; 4 nonprofit or public body with a published conflict policy, an independent board, and external review.
- G.1 Legal form alone. "Nonprofit", "academic", or "public body" with no published conflict policy floors at 2, not 3. Anchor 3 requires the policy.
- G.2 What counts as a conflict policy. A published policy floors at 3 (nonprofit or public body) or 2 (for-profit or PBC) only if it names at least two of: recusal for financial interests, no outcome-contingent fees, no investments or side grants from evaluated labs, disclosure of lab-linked revenue. Statutory civil-service conflict rules count for a public body.
- G.3 Lab-investor money in a round. Participation by a fund that has invested in a frontier developer caps at 2 even with mission board seats.
- G.4 Member-governed consortia whose members include the labs being benchmarked cap at 2.
- G.5 Political steerability. A public body whose leadership or mandate changed more than once in twelve months caps at 3.
- G.6 Ownership at 20% or more by a lab or a lab investor is 0 (same threshold as F.1).
- G.7 Four requires tier-1 evidence of an independent board and of external review.
3. Personnel (P)
Anchors: 0 leaders hold governance roles at an evaluated lab with no recusal; 1 leaders hold equity or advisory roles at labs, informal recusal; 2 frequent two-way hiring, recusal on request; 3 recusal policy and disclosure of lab ties; 4 cooling-off periods, disclosed ties, no equity in labs.
- P.1 Lab board or committee seats. A founder, executive, or director holding a board or committee seat at an evaluated lab caps at 1 with a personal recusal and at 0 without one.
- P.2 Advisory roles. An executive or director advising an evaluated lab caps at 1.
- P.3 Two-way hiring. Documented hiring from, or departures to, evaluated labs without a cooling-off rule caps at 2. An organizational cooling-off policy or a statutory post-employment rule for the body's staff lifts the cap.
- P.4 Seats one hop from a lab. A board member or advisor who is a partner at a lab investor, a founder of a lab contractor, or an employee of a lab caps at 2 regardless of recusal.
- P.5 Recusal policy. A published mandatory recusal rule plus disclosure of individual lab ties floors at 3.
- P.6 Absence of evidence. "No lab roles found" floors at 2, not 3. A 3 needs the policy in P.5.
- P.7 Employees of a lab (a first-party team) or leadership that moved to a lab under an ownership transaction: 0 or 1 by P.1.
- P.8 Four requires a cooling-off rule, disclosed ties, and no equity, with tier-1 evidence (C10).
4. Access depth, lab-granted (A)
Anchors: 0 public API only; 1 pre-release API with safeguards on; 2 pre-release with safeguards off or extended time; 3 helpful-only or weights-level access, chain of thought, logs, on-site; 4 embedded, training-time, or incident access.
- A.1 Score what was granted. The deepest documented engagement sets the floor. Access is recorded as revocable unless A.2 applies; the mechanism tag on the card says so.
- A.2 Enforceable access. Access resting on statute, regulation, or a court-enforceable agreement floors at 3 even if not yet exercised, and the card tags it statutory.
- A.3 Voluntary access. "The lab can decline to renew" is recorded as a mechanism, not as a cap; it does not lower the value by itself.
- A.4 Government agreements (memoranda, testing agreements, classified work): 3.
- A.5 Embedded, on-site, or incident access: 4.
- A.6 No engagement yet: 0.
5. Scope control (S)
Anchors: 0 lab defines tasks and can decline findings; 1 lab defines scope, evaluator picks methods; 2 scope negotiated per engagement; 3 evaluator sets scope and can add questions; 4 evaluator sets scope, can investigate incidents, can refuse sign-off.
- S.1 Lab-set window or excluded questions in the most recent major engagement caps at 3.
- S.2 Cannot add questions: cap 2.
- S.3 Regulator as client setting scope by tender or contract: 3.
- S.4 Statutory scope. A public body with the legal power to define scope and compel participation: 4. Voluntary memoranda: 3.
- S.5 Four needs incident or sign-off rights. Self-directed research with no engagement to refuse floors at 3. A 4 requires documented incident-investigation rights or a documented right to refuse sign-off.
- S.6 Co-authoring research with a lab while evaluating it caps at 3.
6. Publication rights (R)
Anchors: 0 no publication, or lab approval required; 1 lab-edited summaries only; 2 publishes, lab reviews with broad redaction; 3 publishes, redaction limited to security, redactions disclosed; 4 full editorial control, record of adverse findings, redaction statements.
- R.1 Security-only redactions with a public statement of what was redacted floors at 3.
- R.2 Editorial feedback. Lab feedback on emphasis, structure, or tone that the evaluator incorporated caps at 3 if the evaluator discloses that it happened and what changed, and at 2 if it does not.
- R.3 System cards only. Findings that reach the public only as lab-summarized system-card citations cap at 1; if the organization also publishes its own reports on other work, cap 2.
- R.4 Statutory non-publication. A public body is scored on what reaches the public: nothing per model, 1; aggregated trends, 2; reports to a regulator only, 2. The card tags the mechanism statutory so it is not read as a lab holding the pen.
- R.5 Four. No lab pre-publication review, plus at least one published finding the evaluated lab did not welcome (a safety failure, a disputed result, a low grade the lab contested). A leaderboard ranking alone is not an adverse finding. Track record is load-bearing here.
- R.6 Funder approval over disclosure. A contract or arrangement under which a lab controlled whether or when the evaluator could disclose who funded a benchmark or evaluation caps at 2 for the period it applied and at 3 after the evaluator published a disclosure policy.
- R.7 "Publishes" alone. A statement that an organization publishes its results, with no statement about who reviews them before publication, floors at 2. A 3 needs either no lab review or security-only redactions disclosed (R.1).
7. Method transparency (M)
Anchors: 0 closed; 1 summaries only; 2 methods described in prose; 3 tasks or code partly open; 4 open code, tasks, reproducible runs, factsheets.
- M.1 Four requires open code and open tasks and evidence that a third party has re-run them or that per-evaluation factsheets are published. Open code alone is 3.
- M.2 Prose-only methods: 2. Summaries: 1.
- M.3 A public leaderboard with open tasks and logs: 3; with third-party re-runs documented: 4.
8. Role incompatibility (X)
Anchors: 0 sells defense or monitoring products to evaluated labs; 1 sells products or services other than the evaluation itself to labs or to their customers; 2 consults for labs; 3 tools are open or free to the ecosystem; 4 no commercial products.
- X.1 Selling the fix. Guardrails, monitoring, or defense products sold to an evaluated lab: 0. A product sold to unspecified parties in the evaluated ecosystem caps at 2 until a lab customer is documented, and the card carries a document request.
- X.2 Products and services to labs. Data, tooling, licensed evaluation products, or deployment services delivered to an evaluated lab for money: 1. The fee for evaluating the lab's own model is not a second role; it is scored on funding (F.3) and does not cap here.
- X.3 Consulting. Paid consultation or framework work for a lab: 2. Paid consultation clients listed on a transparency page count.
- X.4 Free or open tools used by labs: 3. Labs using a free tool are not customers.
- X.5 Financial interest. An evaluator that invests its own funds in evaluated labs caps at 2; in a diversified portfolio that may include listed labs, 3.
- X.6 Co-producing a benchmark with a vendor owned by a lab caps at 3.
- X.7 Four requires a bounded statement of no commercial products with a span and a second source (C16).
- X.8 Undisclosed earned revenue. Program-service or contract revenue in a filing whose clients are not disclosed caps at 3 until the clients are named; if a lab is named among them, X.2 applies.
9. Track record and mis-dimensioned facts
A pledge to evaluate is not an evaluation. An organization with no published evaluation of a frontier model floors at 0 on access and takes no floor above 2 on scope, publication, or methods. This does not apply to a public body whose access or scope rests on statute or regulation: a legal power is not a pledge, and A.2 and S.4 govern those dimensions for it.
A signal whose fact belongs to another dimension (an ownership fact recorded on access, a funding-disclosure fact recorded on methods) is informational on the dimension it sits on (bound: null, bound_note naming this section) and is re-recorded on the right dimension as a new signal.
10. Role classification
Role is derived, not typed. python -m bench verify rejects a stored role that disagrees.
- Government: a public body (type
gov). - First-party: a unit of a frontier developer (type
bigtech). - Vendor: venture-backed (type
vc), or a role-incompatibility value of 1 or 0 (sells products or services to labs), or a private company that consults for labs (X of 2 with typeprivate). - Independent referee: nonprofit, public benefit corporation, or academic group whose published work is evaluation or research, with no products sold to evaluated labs. Lab fees for the evaluation itself do not change the class; they change the funding value.
- Benchmark or consortium: an organization whose evaluation output is a leaderboard or a standard rather than engagements keeps
role: benchmark; it is placed in a list by its owner's type (academic and nonprofit with referees; consortia funded by member labs with commercial and first-party). - Expected entrant: watchlist only, never ranked.
List placement on the site: independent referees; government institutes; commercial and first-party (vendors, lab units, lab-funded consortia).
11. Population
Scope (D-001): an organization is in scope if it evaluates or red-teams frontier models for safety-relevant properties, meaning dangerous capabilities (biological, chemical, cyber, autonomy), misuse, alignment and scheming, security, safeguards, and incident investigation, or if it is a public body or standards body with a role in that layer. Capability and performance leaderboards, and organizations whose only frontier work is capability benchmarking, are out of scope; if already in the directory they are retained with status: out-of-scope, unranked, shown in their own section, and excluded from every statistic.
Inclusion, within that scope: cited as an external evaluator or red team in at least one frontier system card or government evaluation report in the last 24 months; or named in a statute, code, or standard as an evaluator; or operating a safety benchmark cited by a frontier developer for a frontier model. data/exclusions.json records every candidate checked, the criterion applied, and the result. A hypothetical composite is never scored.
12. People
Public roles only. No inference about motive or timing. Every named individual receives their card and a reply window, the same as an organization (bench outreach --people). An open question that concerns a named person's gift or role is phrased as a request for a document, never as an unresolved suspicion, and goes to the person before it is published.
13. Weights
Each preset in data/presets.json carries a derivation paragraph naming the persona and the external instrument the weights follow. The lab-procurement preset was fixed in the seed script on 15 September 2026 before the population pass of the same day; it was not registered anywhere outside this repository, so it is called the confirmatory preset, not a pre-registered one. The other presets are sensitivity checks.
Glossary
- Hop. Distance from a frontier lab in the ledger: hop 0 is a lab; hop 1 a lab investor, observer, board member, employee, founder, or contractor; higher hops run through principals or funders.
- Signal. One dated claim, for or against an organization, on one dimension, citing at least one source, with a bound that names the anchor it supports.
- Bound. What a signal does to a value: an against-signal caps it, a for-signal floors it. The value is the tightest cap or, failing any cap, the highest floor.
- Evidence policy. The rule for which signals count: leads included, standard (the default: at least one confirmed source), against interest, verified spans, primary only.
- Band. Independence has floors. An evidenced 0 on a conflict dimension (funding, governance, personnel, role incompatibility, scope, publication) is a disqualifying floor; a 1 on one of those a conditional floor; otherwise clear. Access and methods never set a band. The number ranks within the band and is hidden until you choose weights.
- Unevidenced. No admissible signal sets a bound on that dimension under the chosen policy. Shown as a dash; excluded from the score; counted in coverage.
- Held. A bound that the evidence rules would not let stand at its declared value: an extreme without a quoted span, a 4 without a second source, or a bound from sources that were not all confirmed.
- Checked and not found (bounded negative). A recorded search that found nothing, naming the corpus, the snapshot date, and the query. It supports 'not in that corpus on that date' and nothing more.
- Quarantined. A ledger row or source whose re-fetch disagreed with the recorded figure, or that could not be verified. It stays visible with both values and counts for nothing.
- Imported. A row copied from another project's ledger and not yet re-derived from its source. A lead, not evidence.
- Confirmed. Re-fetched from the cited source by a named person on a date. A span is confirmed when the quoted text is found verbatim on the page.
- Mechanism. Why an access or publication value is what it is: statutory, lab-controlled, self-imposed, or unknown.
- Dissent. On every card, the strongest case that the weakest dimension should be one notch lower, and one notch higher.
- Out of scope, retained. An organization outside the scope in RULES 11, kept with its records unchanged, unranked, and excluded from every statistic (decision D-001).