Cases that set the bar
Six cases from the last two years, read for what they reveal about each dimension: five engagements and one investor overlap. Epoch AI is retained out of scope as a capability benchmark (decision D-001); its FrontierMath case stays because it set the funder-disclosure norm the safety evaluators now follow. Longer treatments and sources are in the repo under paper/.
METR and Redwood inside OpenAI
Two METR staff and a Redwood contractor spent six days on OpenAI's premises investigating how roughly 700 agents coordinated a hack of Hugging Face. They took no payment, wrote their own report, and published a statement on what was redacted. OpenAI set the window (June 26 to July 13), could redact any non-public information, and gave feedback on emphasis and tone. The later breach of OpenAI's own research cluster fell outside scope. This is what the access and publication dimensions look like at their current best, and what the scope dimension looks like when the audited party still holds the pen.
Sources: METR, "Brief independent investigation of agents' behavior in the OpenAI / Hugging Face incident," 26 Aug 2026; critique of the scope in the repo under paper/.
Epoch AI and FrontierMath
OpenAI commissioned 300 problems, owned them, and saw the statements and solutions; a contract barred Epoch from saying so until the day o3's score was announced. More than sixty contributing mathematicians were not told. Epoch acknowledged the communication failure and built a 50-problem holdout OpenAI cannot see. Publication rights include the right to disclose who paid.
Sources: Epoch AI, "Clarifying the creation and use of the FrontierMath benchmark," 23 Jan 2025.
SecureBio and the OpenAI Foundation
SecureBio accepted a foundation grant for its Detection division while its AI division evaluates OpenAI models under contracts OpenAI pays for. Leadership disclosed the arrangement, described the firewall, and committed to resign and go public if the grant were ever used as leverage. The foundation's endowment is a stake in OpenAI and the boards largely overlap; the disclosure is what separates this from the Epoch case.
Sources: SecureBio's own disclosure of the grant (X, 2026).
Gray Swan and OpenAI's safety committee
Gray Swan's co-founder and chief scientist chairs OpenAI's Safety and Security Committee and sits on its nonprofit board, and recuses himself from Gray Swan's dealings with OpenAI. Gray Swan also sells jailbreak defenses to the labs it red-teams. Recusal handles the individual; it does not address a firm that grades models and sells the remedy.
Sources: OpenAI board announcements (chair role); Forbes, 2024 (recusal quote).
Irregular, Sequoia and the rebrand
Pattern Labs became Irregular in September 2025 with $80M led by Sequoia and Redpoint at a $450M valuation. Sequoia is a significant OpenAI investor, and Irregular's evaluations sit in OpenAI's system cards. Nothing here is improper; it is the plain case of a venture-backed evaluator whose backers profit when the evaluated lab's models ship.
Sources: TechCrunch / company announcement, 17 Sep 2025.
AVERI's double-blind pilot
AVERI, DeepMind, OpenMined and MLCommons ran Gemini 2.5 Flash-Lite against never-seen AILuminate prompts inside a secure enclave, so the lab did not see the test items and the evaluator did not see model internals it was not cleared for. Small model, voluntary, and lab-selected, but it demonstrates that the usual excuse for shallow access has an engineering solution.
Sources: AVERI Pilot Report, Aug 2026.