Holistic Agent Leaderboard (Princeton)
Scorecard
Academic agent leaderboard with fully open methodology; AEF founding member.
Academic, Princeton, US. Benchmark or consortium; listed with independent referees. Confidence med. Domains: benchmarks, autonomy.
8/8 evidenced under the standard policy Scores by preset: Lab procurement 56; Regulator or auditor 55; Public trust 64; Equal weights 59. Weakest evidenced dimension: access depth (lab-granted) 0/4. At least one evidenced dimension at 0: not a candidate for an independence-critical role until the floor moves.
What would move the score. Nothing on independence; access is the gap.
- Funding 3 to 4: score 60
- Governance 2 to 3: score 57
- Personnel 2 to 3: score 58
Dissent, lower. Access is already 0 and cannot go lower, so the weakest movable dimension is Personnel, which should be 1. A co-author is listed at xAI, staff of Anthropic, Google DeepMind and UK AISI are acknowledged on the page, OpenAI and Google supply the API credits, and no recusal or contributor conflict rule exists; lab-affiliated authors on a leaderboard that grades their employers is anchor 1's informal arrangement.
Dissent, higher. Personnel should be 3. Every co-author's affiliation is disclosed on the page, the project sits inside Princeton, whose faculty are bound by a university conflict-of-interest policy, no HAL leader holds a lab role, and the leaderboard's scoring is open code anyone can re-run. Disclosed ties under an institutional policy is anchor 3, not anchor 2.
Each value is the tightest admissible cap or, with no cap, the highest admissible floor, under RULES.md. The binding signal is highlighted. A dash means no admissible signal sets a bound under the standard policy.
- No lab money; academic. floors at 4 by F.11
Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed2026-09-14 - University-hosted leaderboard; no lab money found, and no bounded negative on file (no filing or index searched yet). caps at 3 by F.11binding“HAL is funded by Coefficient Giving , Schmidt Sciences , the Princeton AI Lab”
Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed2026-09-14 - HAL is funded by Coefficient Giving, whose principal is an Anthropic investor, and receives API credits from OpenAI and Google to evaluate their models. caps at 3 by F.6binding“We are grateful to OpenAI and Google for providing API credits to evaluate their models”
Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed Anthropic raises $124 million Series A (Anthropic, 2021-05-28) tier 3 self (anthropic)confirmed2026-09-15 - Document request: Request to the HAL team: the amount or approximate value of the OpenAI and Google API credits received and whether any evaluation depends on them, plus a funder list with amounts so a bounded negative on lab money can be recorded.
- University-hosted; AEF member. floors at 2 by G.1binding“By the SAgE team at Princeton University”
Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed AI Evaluator Forum launch and AEF-1 (AI Evaluator Forum, 2025-12-04) tier 3 self (aef)confirmed2026-09-14
- Academic staff. floors at 2 by P.6“Sayash Kapoor Princeton University Benedikt Stroebl Princeton University”
Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed2026-09-14 - HAL's author list includes a co-author affiliated with xAI, and its acknowledgments name staff at Anthropic, Google DeepMind and UK AISI; no contributor recusal or cooling-off rule is published. caps at 2 by P.4binding“Yifei Zhou xAI”
Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed2026-09-15 - Document request: Request to the HAL team at Princeton: the contributor conflict-of-interest or recusal rule that applies to the leaderboard, if any, and the date on which the xAI-affiliated co-author's affiliation began relative to the runs they contributed.
- Public API access only. caps at 0 by A.6binding“We are grateful to OpenAI and Google for providing API credits to evaluate their models”
Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed2026-09-14
- Own benchmark design. floors at 3 by S.5binding“We have paused updating HAL leaderboard with new models and are currently focusing on measuring reliability in AI agents”
Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed2026-09-14
- Results published without review. floors at 3 by R.5binding“The standardized, cost-aware, and third-party leaderboard for evaluating agents”
Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed2026-09-14
- Open code and logs. floors at 3 by M.1binding“HAL is an open-source project and we welcome contributions from the community”
Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed2026-09-14
- No products. floors at 4 by X.7binding“HAL is an open-source project”
Holistic Agent Leaderboard (Princeton, 2026) tier 3 self (hal)confirmed2026-09-14
Values under each evidence policy
| Dimension | Leads included | Standard | Against interest | Verified spans | Primary only |
|---|---|---|---|---|---|
| Funding | 3 | 3 | 3 | 3 | – |
| Governance | 2 | 2 | 2 | 2 | – |
| Personnel | 2 | 2 | 2 | 2 | – |
| Access depth (lab-granted) | 0 | 0 | 0 | 0 | – |
| Scope control | 3 | 3 | – | 3 | – |
| Publication rights | 3 | 3 | – | 3 | – |
| Method transparency | 3 | 3 | – | 3 | – |
| Role incompatibility | 3 | 3 | – | 3 | – |
| Score, Lab procurement | 56 Disqualifying floor, 8/8 | 56 Disqualifying floor, 8/8 | 40 Disqualifying floor, 4/8 | 56 Disqualifying floor, 8/8 | – Unevidenced, 0/8 |
| Score, Regulator or auditor | 55 Disqualifying floor, 8/8 | 55 Disqualifying floor, 8/8 | 39 Disqualifying floor, 4/8 | 55 Disqualifying floor, 8/8 | – Unevidenced, 0/8 |
| Score, Public trust | 64 Disqualifying floor, 8/8 | 64 Disqualifying floor, 8/8 | 56 Disqualifying floor, 4/8 | 64 Disqualifying floor, 8/8 | – Unevidenced, 0/8 |
| Score, Equal weights | 59 Disqualifying floor, 8/8 | 59 Disqualifying floor, 8/8 | 44 Disqualifying floor, 4/8 | 59 Disqualifying floor, 8/8 | – Unevidenced, 0/8 |
Traced money and ties
no inflow rows yet. Confirmed rows: 0 of 0. Dollar sums: confirmed + unaudited USD rows only; imported figures are not re-derived and never summed. Second hop traced for 0 of 0 sources. Ties: none within two steps recorded. Bounded negatives on file: 0.
Ledger
No rows yet. Add one in data/ledger/.
Funding graph, focused
Neighbours at full strength, everything else faded. Hover any name to move the focus; click a name to open its page.