Evaluator Bench

The rubric and the rules

Eight dimensions, each scored 0 to 4 from public evidence, then weighted. The dimensions follow the AI Evaluator Forum's AEF-1 operating conditions and the financial-audit independence rules that Illinois SB 315 imports for frontier AI, with two additions the field tends to skip: who owns the evaluator, and whether it sells fixes to the companies it grades. A value is not a curator's impression: each signal names the anchor it supports and the rule below that says so, and the value is the tightest cap or, with no cap, the highest floor.

Funding

Where the money comes from, and whether any of it comes from the developers being evaluated or their investors.

  • 0Owned or controlled by a frontier developer or a lab investor: a stake of 20% or more, or a business unit.
  • 1Material revenue or investment from evaluated labs or their investors.
  • 2Labs pay per engagement; otherwise diversified.
  • 3Mostly philanthropic or public money; some lab-linked pooled funds.
  • 4No lab money; diversified philanthropic or public funding, disclosed.

Governance

Legal form, board, and whether a conflict-of-interest policy is published.

  • 0Unit or subsidiary of a lab or a lab's investor.
  • 1VC-backed for-profit with no published COI policy.
  • 2For-profit or PBC with a published COI policy.
  • 3Nonprofit or public body with a COI policy.
  • 4Nonprofit or public body, published COI policy, independent board, external review.

Personnel

Board seats, equity, advisory roles, and the revolving door between evaluator and lab.

  • 0Leaders hold governance roles at an evaluated lab, no recusal.
  • 1Leaders hold equity or advisory roles at labs; informal recusal.
  • 2Frequent two-way hiring; recusal on request.
  • 3Recusal policy and disclosure of lab ties.
  • 4Cooling-off periods, disclosed ties, no equity in labs.

Access depth (lab-granted)

The deepest access labs have actually granted this evaluator in practice: public API, pre-release API, safeguards off, weights and logs, or embedded during training. This is granted access, not an institutional right or a capability measure: a low score can mean labs did not grant access, not that the evaluator lacks competence. It records who labs chose to let in.

  • 0Public API only.
  • 1Pre-release API with safeguards on.
  • 2Pre-release with safeguards off or extended time.
  • 3Helpful-only or weights-level access, chain of thought, logs, on-site.
  • 4Embedded, training-time, or incident access.

Scope control

Who decides what gets tested, for how long, and whether the evaluator can refuse to sign off.

  • 0Lab defines tasks and can decline findings.
  • 1Lab defines scope; evaluator picks methods.
  • 2Scope negotiated per engagement.
  • 3Evaluator sets scope and can add questions.
  • 4Evaluator sets scope, can investigate incidents, can refuse sign-off.

Publication rights

Whether findings reach the public unedited, and whether adverse findings have been published.

  • 0No publication, or lab approval required.
  • 1Lab-edited summaries only.
  • 2Publishes; lab reviews with broad redaction.
  • 3Publishes; redaction limited to security; redactions disclosed.
  • 4Full editorial control, record of adverse findings, redaction statements.

Method transparency

Whether evaluation code, tasks, and conditions are open enough for others to reproduce.

  • 0Closed.
  • 1Summaries only.
  • 2Methods described in prose.
  • 3Tasks or code partly open.
  • 4Open code, tasks, reproducible runs, factsheets.

Role incompatibility

Whether the organization both grades labs and sells to them — the audit-plus-consulting problem. Selling products or services to an evaluated lab is a role conflict, not proof that any given evaluation was wrong; a vendor can be a competent tester and still be structurally compromised as a referee.

  • 0Sells defense or monitoring products to evaluated labs.
  • 1Sells products or services other than the evaluation itself to labs or to their customers.
  • 2Consults for labs.
  • 3Tools are open or free to the ecosystem.
  • 4No commercial products.

What raises and lowers a score

The list is deliberately concrete: each item is something you can verify from a filing, a contract term, a system card, or a published policy.

Raises the score

  • A written policy refusing money from frontier developers, including donations directed by their staff.
  • A published conflict-of-interest policy: no outcome-contingent fees, no side grants or investments from labs being evaluated, mandatory recusal for anyone with a financial interest.
  • No single funder above a stated share of budget; funders disclosed by name.
  • Publication rights fixed in the contract before work starts, with redaction limited to security-sensitive or privileged material and a public statement of what was redacted.
  • A record of publishing findings the lab did not like: scheming, reward hacking, shutdown resistance, safeguard failures.
  • Open evaluation code and task suites so results can be re-run by others.
  • On-site, weights-level, or training-time access rather than a few weeks on an API.
  • Access that does not depend on the lab's goodwill: statute, regulation, or a court-enforceable agreement.
  • Membership in a standards body (AEF-1, AVERI pilots) and disclosure of operating conditions for each evaluation.
  • Cooling-off periods before staff move to labs; equity in labs disclosed or prohibited.
  • A client base spread across many developers, so no lab is a dominant revenue source.
  • Double-blind or secure-enclave protocols that let the evaluator work without the lab seeing the test items.

Lowers the score

  • The lab pays for the evaluation of its own model. Common, and the single most under-discussed conflict in the field.
  • Investors shared with the labs: a venture firm that backs both the evaluator and OpenAI, or a chip vendor buying the evaluator while anchoring a lab's IPO.
  • A stake of 20% or more held by a frontier developer or a lab investor, or status as a business unit of one.
  • Leadership holding a board seat, safety-committee chair, or advisory role at a lab the organization evaluates, even with recusal.
  • Selling defenses, guardrails, or monitoring products to the same labs it evaluates: the audit-plus-consulting problem that Sarbanes-Oxley separated in accounting.
  • A contract that forbids disclosing who funded a benchmark or evaluation.
  • The lab sets the scope, the time window, and which questions are out of bounds.
  • The lab reviews drafts for tone and emphasis, not only for security redactions.
  • Only aggregated or lab-summarized results reach the public.
  • Access that the lab can decline to renew with no consequence.
  • Political or budgetary dependence that can redirect a government institute's mandate within a year.
  • Co-authoring research with a lab while also serving as its external evaluator.
  • Free tokens and compute from the lab, when they are a material share of operating capacity.
  • Heavy talent flow in both directions between the evaluator and the labs.
  • No track record: a pledge to evaluate is not an evaluation.

Weights

The rules

RULES: how a signal becomes a value

Version 0.1, 15 September 2026. These rules are written before they are applied and are applied to every organization the same way. A value on a dimension is not a curator's impression of the file; it is the tightest admissible cap (from evidence against) or, failing any cap, the highest admissible floor (from evidence for), where every signal names the anchor it supports under the rule that says so. Where a floor and a cap disagree, the assessment carries a written resolution naming the rule that decides it, and the disagreement stays on the card.

python -m bench verify enforces the mechanics (CONTRACT C22 to C25). This file is the meaning. Change the rule before changing a value.

0. Mechanics

1. Funding (F)

Anchors: 0 owned or controlled by a frontier developer or a lab investor (a stake of 20% or more, or a business unit); 1 material revenue or investment from evaluated labs or their investors; 2 labs pay per engagement, otherwise diversified; 3 mostly philanthropic or public money, some lab-linked pooled funds; 4 no lab money, diversified philanthropic or public funding, disclosed and bounded.

2. Governance (G)

Anchors: 0 unit or subsidiary of a lab or a lab's investor; 1 venture-backed for-profit with no published conflict-of-interest policy; 2 for-profit or public benefit corporation with a published conflict policy; 3 nonprofit or public body with a conflict policy; 4 nonprofit or public body with a published conflict policy, an independent board, and external review.

3. Personnel (P)

Anchors: 0 leaders hold governance roles at an evaluated lab with no recusal; 1 leaders hold equity or advisory roles at labs, informal recusal; 2 frequent two-way hiring, recusal on request; 3 recusal policy and disclosure of lab ties; 4 cooling-off periods, disclosed ties, no equity in labs.

4. Access depth, lab-granted (A)

Anchors: 0 public API only; 1 pre-release API with safeguards on; 2 pre-release with safeguards off or extended time; 3 helpful-only or weights-level access, chain of thought, logs, on-site; 4 embedded, training-time, or incident access.

5. Scope control (S)

Anchors: 0 lab defines tasks and can decline findings; 1 lab defines scope, evaluator picks methods; 2 scope negotiated per engagement; 3 evaluator sets scope and can add questions; 4 evaluator sets scope, can investigate incidents, can refuse sign-off.

6. Publication rights (R)

Anchors: 0 no publication, or lab approval required; 1 lab-edited summaries only; 2 publishes, lab reviews with broad redaction; 3 publishes, redaction limited to security, redactions disclosed; 4 full editorial control, record of adverse findings, redaction statements.

7. Method transparency (M)

Anchors: 0 closed; 1 summaries only; 2 methods described in prose; 3 tasks or code partly open; 4 open code, tasks, reproducible runs, factsheets.

8. Role incompatibility (X)

Anchors: 0 sells defense or monitoring products to evaluated labs; 1 sells products or services other than the evaluation itself to labs or to their customers; 2 consults for labs; 3 tools are open or free to the ecosystem; 4 no commercial products.

9. Track record and mis-dimensioned facts

A pledge to evaluate is not an evaluation. An organization with no published evaluation of a frontier model floors at 0 on access and takes no floor above 2 on scope, publication, or methods. This does not apply to a public body whose access or scope rests on statute or regulation: a legal power is not a pledge, and A.2 and S.4 govern those dimensions for it.

A signal whose fact belongs to another dimension (an ownership fact recorded on access, a funding-disclosure fact recorded on methods) is informational on the dimension it sits on (bound: null, bound_note naming this section) and is re-recorded on the right dimension as a new signal.

10. Role classification

Role is derived, not typed. python -m bench verify rejects a stored role that disagrees.

List placement on the site: independent referees; government institutes; commercial and first-party (vendors, lab units, lab-funded consortia).

11. Population

Scope (D-001): an organization is in scope if it evaluates or red-teams frontier models for safety-relevant properties, meaning dangerous capabilities (biological, chemical, cyber, autonomy), misuse, alignment and scheming, security, safeguards, and incident investigation, or if it is a public body or standards body with a role in that layer. Capability and performance leaderboards, and organizations whose only frontier work is capability benchmarking, are out of scope; if already in the directory they are retained with status: out-of-scope, unranked, shown in their own section, and excluded from every statistic.

Inclusion, within that scope: cited as an external evaluator or red team in at least one frontier system card or government evaluation report in the last 24 months; or named in a statute, code, or standard as an evaluator; or operating a safety benchmark cited by a frontier developer for a frontier model. data/exclusions.json records every candidate checked, the criterion applied, and the result. A hypothetical composite is never scored.

12. People

Public roles only. No inference about motive or timing. Every named individual receives their card and a reply window, the same as an organization (bench outreach --people). An open question that concerns a named person's gift or role is phrased as a request for a document, never as an unresolved suspicion, and goes to the person before it is published.

13. Weights

Each preset in data/presets.json carries a derivation paragraph naming the persona and the external instrument the weights follow. The lab-procurement preset was fixed in the seed script on 15 September 2026 before the population pass of the same day; it was not registered anywhere outside this repository, so it is called the confirmatory preset, not a pre-registered one. The other presets are sensitivity checks.

Glossary