Evaluator Bench
CAIS, 2026; retrieved 2026-09-14 tier 3 self (cais)confirmed

Center for AI Safety

https://safe.ai/

WMDP, HLE benchmarks; director advises xAI. Cited from a search snippet; not re-fetched in full. [verify 2026-09-15] fetched; page reachable, citation matches

Signals citing this source

  • Center for AI Safety, P (against): Director is a safety adviser to xAI while the organization's benchmarks are used to grade xAI models.
  • Center for AI Safety, M (for): Open benchmarks widely used.“CAIS develops foundational benchmarks and methods which concretize the problem”
  • Center for AI Safety, X (against): Co-produced Humanity's Last Exam with Scale AI, a Meta-owned vendor.
  • Center for AI Safety, G (for): Nonprofit.“is a San Francisco-based research and field-building nonprofit”
  • Center for AI Safety, A (against): Public-model access.
  • Center for AI Safety, S (for): Sets own benchmark design.“CAIS develops foundational benchmarks and methods which concretize the problem”
  • Center for AI Safety, R (for): Publishes results.“The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems Mar 5, 2025”

Ledger rows citing this URL

  • R43 principal at Center for AI Safety confirmed