Evaluator Bench
evaluator; distance from a lab: unattributed

Holistic Agent Leaderboard (Princeton)

Scorecard

Academic agent leaderboard with fully open methodology; AEF founding member.

Academic, Princeton, US. Benchmark or consortium; listed with independent referees. Confidence med. Domains: benchmarks, autonomy.

Disqualifying floor 8/8 evidenced under the standard policy Scores by preset: Lab procurement 56; Regulator or auditor 55; Public trust 64; Equal weights 59. Weakest evidenced dimension: access depth (lab-granted) 0/4. At least one evidenced dimension at 0: not a candidate for an independence-critical role until the floor moves.

What would move the score. Nothing on independence; access is the gap.

Dissent, lower. Access is already 0 and cannot go lower, so the weakest movable dimension is Personnel, which should be 1. A co-author is listed at xAI, staff of Anthropic, Google DeepMind and UK AISI are acknowledged on the page, OpenAI and Google supply the API credits, and no recusal or contributor conflict rule exists; lab-affiliated authors on a leaderboard that grades their employers is anchor 1's informal arrangement.

Dissent, higher. Personnel should be 3. Every co-author's affiliation is disclosed on the page, the project sits inside Princeton, whose faculty are bound by a university conflict-of-interest policy, no HAL leader holds a lab role, and the leaderboard's scoring is open code anyone can re-run. Disclosed ties under an institutional policy is anchor 3, not anchor 2.

Each value is the tightest admissible cap or, with no cap, the highest admissible floor, under RULES.md. The binding signal is highlighted. A dash means no admissible signal sets a bound under the standard policy.

Funding 3/4
Anchor 3: Mostly philanthropic or public money; some lab-linked pooled funds.
evidence: tier 3 (self)
held at a supportable anchor
Capped at 3 by hal.09 (F.11): University-hosted leaderboard; no lab money found, and no bounded negative on file (no filing or index searched yet).; capped at 3 by hal.10 (F.6): HAL is funded by Coefficient Giving, whose principal is an Anthropic investor, and receives API credits from OpenAI and Google to evaluate their models.. Floors up to 3 from hal.01 do not exceed the cap.
Curator: Anchor 3 pending a bounded negative (university sponsorship disclosures or a funder index search).
Governance 2/4
Anchor 2: For-profit or PBC with a published COI policy.
evidence: tier 3 (self)
Floored at 2 by hal.05 (G.1): University-hosted; AEF member.. No admissible signal caps it.
Curator: Anchor 3: university-hosted academic project and AEF member (hal.05). No tier-1 source for an independent board or external review, so 4 is unearned under the tier rule (C10).
Personnel 2/4
Anchor 2: Frequent two-way hiring; recusal on request.
evidence: tier 3 (self)
Capped at 2 by hal.11 (P.4): HAL's author list includes a co-author affiliated with xAI, and its acknowledgments name staff at Anthropic, Google DeepMind and UK AISI; no contributor recusal or cooling-off rule is published.. Floors up to 2 from hal.06 do not exceed the cap.
Curator: Anchor 3 (provisional): academic staff (hal.06); no lab-tie disclosures or cooling-off periods are documented at any tier, so 4 is unearned under the tier rule (C10). Provisional readings are open to public correction through the contribution path.
  • Academic staff. floors at 2 by P.6“Sayash Kapoor Princeton University Benedikt Stroebl Princeton University”
    Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed2026-09-14
  • HAL's author list includes a co-author affiliated with xAI, and its acknowledgments name staff at Anthropic, Google DeepMind and UK AISI; no contributor recusal or cooling-off rule is published. caps at 2 by P.4binding“Yifei Zhou xAI”
    Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed2026-09-15
  • Document request: Request to the HAL team at Princeton: the contributor conflict-of-interest or recusal rule that applies to the leaderboard, if any, and the date on which the xAI-affiliated co-author's affiliation began relative to the runs they contributed.
Access depth (lab-granted) 0/4
Anchor 0: Public API only.
evidence: tier 3 (self)
mechanism: lab-controlled
Capped at 0 by hal.04 (A.6): Public API access only..
Scope control 3/4
Anchor 3: Evaluator sets scope and can add questions.
evidence: tier 3 (self)
Floored at 3 by hal.07 (S.5): Own benchmark design.. No admissible signal caps it.
Publication rights 3/4
Anchor 3: Publishes; redaction limited to security; redactions disclosed.
evidence: tier 3 (self)
mechanism: self-imposed
Floored at 3 by hal.03 (R.5): Results published without review.. No admissible signal caps it.
Method transparency 3/4
Anchor 3: Tasks or code partly open.
evidence: tier 3 (self)
Floored at 3 by hal.02 (M.1): Open code and logs.. No admissible signal caps it.
Role incompatibility 3/4
Anchor 3: Tools are open or free to the ecosystem.
evidence: tier 3 (self)
held at a supportable anchor
Floored at 3 by hal.08 (X.7): No products.. No admissible signal caps it. Held: C16: a 4 needs a tier-1/2 source or two independent sources, one not self-published.

Values under each evidence policy

DimensionLeads includedStandardAgainst interestVerified spansPrimary only
Funding3333
Governance2222
Personnel2222
Access depth (lab-granted)0000
Scope control333
Publication rights333
Method transparency333
Role incompatibility333
Score, Lab procurement56 Disqualifying floor, 8/856 Disqualifying floor, 8/840 Disqualifying floor, 4/856 Disqualifying floor, 8/8Unevidenced, 0/8
Score, Regulator or auditor55 Disqualifying floor, 8/855 Disqualifying floor, 8/839 Disqualifying floor, 4/855 Disqualifying floor, 8/8Unevidenced, 0/8
Score, Public trust64 Disqualifying floor, 8/864 Disqualifying floor, 8/856 Disqualifying floor, 4/864 Disqualifying floor, 8/8Unevidenced, 0/8
Score, Equal weights59 Disqualifying floor, 8/859 Disqualifying floor, 8/844 Disqualifying floor, 4/859 Disqualifying floor, 8/8Unevidenced, 0/8

Traced money and ties

no inflow rows yet. Confirmed rows: 0 of 0. Dollar sums: confirmed + unaudited USD rows only; imported figures are not re-derived and never summed. Second hop traced for 0 of 0 sources. Ties: none within two steps recorded. Bounded negatives on file: 0.

Ledger

No rows yet. Add one in data/ledger/.

Funding graph, focused

Neighbours at full strength, everything else faded. Hover any name to move the focus; click a name to open its page.

The funding graph around Holistic Agent Leaderboard (Princeton) Teal lines are money (transfers); amber lines are roles. Solid: confirmed; dashed: imported. Hover a name to isolate it and its neighbours; hover a line for the row. Every element is a row in data/ledger/. Labs Direct lab ties Two steps Three or more Public or unattributed Evaluators Amazon Anthropic G42 Google / Google DeepMind Meta Microsoft OpenAI Thinking Machines Lab xAI AI Safety Fund (Frontier Model Alexandr Wang Andreessen Horowitz Anthropic employees (personal Ben Mann D.E. Shaw Ventures David Farhi Dustin Moskovitz Frontier-lab employees and alu Holden Karnofsky Jaan Tallinn Jeffrey Ladish Macroscopic Ventures (formerly Miles Brundage Neil Chowdhury Nvidia OpenAI employees (personal hol OpenAI Foundation Peter Mattson Sequoia Capital Zico Kolter Adam Gleave Artificial Intelligence Underw Clement Delangue Coefficient Giving (formerly O Good Ventures Foundation Jason Droege Longview Philanthropy Marco Mascorro Paul Christiano Rajiv Dattani Survival and Flourishing Fund Conrad Stosz Dan Hendrycks Mike McCormick Alec Radford Alignment Research Center Constellation Research Center ELMA Philanthropies Emerson Collective EU budget (Digital Europe Prog Fifty Years (50Y) Founders Pledge Founders Pledge (frontier AI f Gates Foundation Halcyon Futures Hillspire (Schmidt family offi Hudson River Trading Jacob Hilton MacKenzie Scott Magarac Venture Partners Obvious Ventures Redpoint Ventures Renaissance Philanthropy Salesforce Ventures Samsung Next Schmidt Sciences Silicon Valley Community Found Skoll Foundation Snowflake Ventures Swish Ventures Sympatico Ventures The Audacious Project (TED) UK government (DSIT) US government (NIST appropriat Valhalla Foundation Vanguard Charitable Wing Venture Capital Y Combinator Andon Labs Apollo Research AVERI Center for AI Safety Epoch AI EquiStamp EU AI Office FAR.AI Gray Swan Hugging Face Open Alignment In Irregular (formerly Pattern La METR Microsoft AI Red Team MLCommons Nemesys Insights Palisade Research RAND Corporation Redwood Research SaferAI Scale AI (SEAL) SecureBio Transluce UK AI Security Institute US CAISI (NIST)