Evaluator Bench
evaluator; distance from a lab: unattributed

Humane Intelligence

Scorecard

Public red-teaming exercises with NIST and IMDA; societal harms rather than catastrophic risk.

Nonprofit, New York, US. Independent referee; listed with independent referees. Confidence med. Domains: misuse, jailbreak.

Conditional floor 8/8 evidenced under the standard policy Scores by preset: Lab procurement 57; Regulator or auditor 57; Public trust 63; Equal weights 56. Weakest evidenced dimension: access depth (lab-granted) 1/4. At least one evidenced dimension at 1: usable with conditions the card names.

What would move the score. Pre-release access for the harms it covers.

Dissent, lower. Access should be 0. No pre-release access is documented anywhere: the public exercises, bias bounties and NIST and IMDA partnerships tested deployed models, no system card names Humane Intelligence as a pre-release tester, and no lab engagement is on record. The card shows 1 only because the anchor-0 claim lacks a quoted span (C15), not because anything above public access was ever granted.

Dissent, higher. Access should be 2. Humane Intelligence co-ran the NIST ARIA pilot and the Singapore IMDA multilingual red-teaming exercise, in which participating developers supplied models under government programs; A.4 treats government testing agreements as 3. If those exercises gave structured access beyond the public API, even with safeguards on, the documented engagement is at least anchor 2.

Each value is the tightest admissible cap or, with no cap, the highest admissible floor, under RULES.md. The binding signal is highlighted. A dash means no admissible signal sets a bound under the standard policy.

Funding 3/4
Anchor 3: Mostly philanthropic or public money; some lab-linked pooled funds.
evidence: tier 3 (self)
Floored at 3 by humane.02 (F.9): Government partners rather than lab clients.. No admissible signal caps it.
Governance 2/4
Anchor 2: For-profit or PBC with a published COI policy.
evidence: tier 3 (self)
Floored at 2 by humane.04 (G.1): Nonprofit.. No admissible signal caps it.
Personnel 2/4
Anchor 2: Frequent two-way hiring; recusal on request.
evidence: tier 3 (self)
Capped at 2 by humane.10 (P.4): Humane Intelligence's advisory group includes a Senior Research Scientist at Meta, and its board president previously worked on OpenAI's Human Data team; no recusal or cooling-off rule is published.. Floors up to 2 from humane.05 do not exceed the cap.
  • No lab roles found. floors at 2 by P.6
    Humane Intelligence (Humane Intelligence, 2026) tier 3 self (humane)confirmed2026-09-14
  • Humane Intelligence's advisory group includes a Senior Research Scientist at Meta, and its board president previously worked on OpenAI's Human Data team; no recusal or cooling-off rule is published. caps at 2 by P.4binding“Dr. Diego Garcia-Olano Senior Research Scientist, Meta”
    Board & Advisory - Humane Intelligence (Humane Intelligence, 2026-07) tier 3 self (humane)confirmed2026-09-15
  • Document request: Request to Humane Intelligence: the board and advisory-group conflict-of-interest policy, and whether advisory-group members employed by a model developer recuse from evaluations involving that developer's models.
  • Document request: Request to Humane Intelligence: whether any cooling-off rule applies to directors previously employed by a model developer, and the dates of the board president's OpenAI employment as stated in the published bio.
Access depth (lab-granted) 1/4
Anchor 1: Pre-release API with safeguards on.
evidence: tier 3 (self)
mechanism: lab-controlled
held at a supportable anchor
Capped at 1 by humane.03 (A.6): Public-model access only.. Held: C15: a 0 needs a quoted span.
Scope control 3/4
Anchor 3: Evaluator sets scope and can add questions.
evidence: tier 3 (self)
Floored at 3 by humane.06 (S.5): Designs its own exercises.. No admissible signal caps it.
Publication rights 3/4
Anchor 3: Publishes; redaction limited to security; redactions disclosed.
evidence: tier 3 (self)
mechanism: self-imposed
Floored at 3 by humane.01 (R.5): Publishes openly; public exercises.. No admissible signal caps it.
Method transparency 2/4
Anchor 2: Methods described in prose.
evidence: tier 3 (self)
Floored at 2 by humane.07 (M.2): Methods published.. No admissible signal caps it.
Role incompatibility 2/4
Anchor 2: Consults for labs.
evidence: tier 3 (self)
held at a supportable anchor
floor and cap disagree
Conflict: floor 3 from humane.08 (X.7) against cap 2 from humane.09 (X.1); resolved at 2 by X.1: humane.08's raw floor of 4 is held to 3 by C15/C16 and still exceeds the X.1 cap; X.1 decides because the paid-services statement is on the organization's own page.
Resolution: X.1 decides 2. humane.08's raw floor of 4 is held to 3 by C15/C16 and still exceeds the X.1 cap; X.1 decides because the paid-services statement is on the organization's own page.

Values under each evidence policy

DimensionLeads includedStandardAgainst interestVerified spansPrimary only
Funding333
Governance222
Personnel2222
Access depth (lab-granted)111
Scope control333
Publication rights33
Method transparency222
Role incompatibility2222
Score, Lab procurement57 Conditional floor, 8/857 Conditional floor, 8/836 Conditional floor, 3/863 Clear, 6/8Unevidenced, 0/8
Score, Regulator or auditor57 Conditional floor, 8/857 Conditional floor, 8/833 Conditional floor, 3/863 Clear, 6/8Unevidenced, 0/8
Score, Public trust63 Conditional floor, 8/863 Conditional floor, 8/845 Conditional floor, 3/862 Clear, 6/8Unevidenced, 0/8
Score, Equal weights56 Conditional floor, 8/856 Conditional floor, 8/842 Conditional floor, 3/858 Clear, 6/8Unevidenced, 0/8

Traced money and ties

no inflow rows yet. Confirmed rows: 0 of 0. Dollar sums: confirmed + unaudited USD rows only; imported figures are not re-derived and never summed. Second hop traced for 0 of 0 sources. Ties: none within two steps recorded. Bounded negatives on file: 0.

Ledger

No rows yet. Add one in data/ledger/.

Funding graph, focused

Neighbours at full strength, everything else faded. Hover any name to move the focus; click a name to open its page.

The funding graph around Humane Intelligence Teal lines are money (transfers); amber lines are roles. Solid: confirmed; dashed: imported. Hover a name to isolate it and its neighbours; hover a line for the row. Every element is a row in data/ledger/. Labs Direct lab ties Two steps Three or more Public or unattributed Evaluators Amazon Anthropic G42 Google / Google DeepMind Meta Microsoft OpenAI Thinking Machines Lab xAI AI Safety Fund (Frontier Model Alexandr Wang Andreessen Horowitz Anthropic employees (personal Ben Mann D.E. Shaw Ventures David Farhi Dustin Moskovitz Frontier-lab employees and alu Holden Karnofsky Jaan Tallinn Jeffrey Ladish Macroscopic Ventures (formerly Miles Brundage Neil Chowdhury Nvidia OpenAI employees (personal hol OpenAI Foundation Peter Mattson Sequoia Capital Zico Kolter Adam Gleave Artificial Intelligence Underw Clement Delangue Coefficient Giving (formerly O Good Ventures Foundation Jason Droege Longview Philanthropy Marco Mascorro Paul Christiano Rajiv Dattani Survival and Flourishing Fund Conrad Stosz Dan Hendrycks Mike McCormick Alec Radford Alignment Research Center Constellation Research Center ELMA Philanthropies Emerson Collective EU budget (Digital Europe Prog Fifty Years (50Y) Founders Pledge Founders Pledge (frontier AI f Gates Foundation Halcyon Futures Hillspire (Schmidt family offi Hudson River Trading Jacob Hilton MacKenzie Scott Magarac Venture Partners Obvious Ventures Redpoint Ventures Renaissance Philanthropy Salesforce Ventures Samsung Next Schmidt Sciences Silicon Valley Community Found Skoll Foundation Snowflake Ventures Swish Ventures Sympatico Ventures The Audacious Project (TED) UK government (DSIT) US government (NIST appropriat Valhalla Foundation Vanguard Charitable Wing Venture Capital Y Combinator Andon Labs Apollo Research AVERI Center for AI Safety Epoch AI EquiStamp EU AI Office FAR.AI Gray Swan Hugging Face Open Alignment In Irregular (formerly Pattern La METR Microsoft AI Red Team MLCommons Nemesys Insights Palisade Research RAND Corporation Redwood Research SaferAI Scale AI (SEAL) SecureBio Transluce UK AI Security Institute US CAISI (NIST)