Evaluator Bench
Apollo Research, 2026; retrieved 2026-09-14 tier 3 self (apollo)confirmed

About

https://www.apolloresearch.ai/about

Ran evaluations for all major labs; partnered with OpenAI on anti-scheming; Watcher product. Cited from a search snippet; not re-fetched in full. [verify 2026-09-15] fetched; page reachable, citation matches

Signals citing this source

  • Apollo Research, R (for): Record of adverse findings published, including in-context scheming and evaluation awareness.“published the first evidence that frontier models can scheme in context”
  • Apollo Research, A (for): Ran pre-deployment evaluations for all major labs; contracted by UK AISI for deception evaluations.“partnered with OpenAI to test their o1 model before public deployment”
  • Apollo Research, S (against): Co-authored anti-scheming research with OpenAI while acting as its external scheming evaluator.“partnered with OpenAI to study anti-scheming interventions on frontier models”
  • Apollo Research, X (against): Sells Watcher, a monitoring product, into the ecosystem it evaluates.“building an AI security tool to monitor frontier AI agents”

Ledger rows citing this URL

  • none