Evaluator Bench
arXiv (Dreadnode authors), 2025-06-17; retrieved 2026-09-15 tier 3 self (dreadnode)confirmed

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

https://arxiv.org/abs/2506.14682

Paper by Dreadnode staff (Dawson, Mulla, Landers, Caldwell); self-published research. Fetched 2026-09-15; span 'challenges from the Crucible challenge environment on the Dreadnode platform' present. Reports results on Claude 3.7 Sonnet, Gemini 2.5 Pro, GPT-4.5 and others: a published evaluation of frontier models.

Signals citing this source

  • Dreadnode, R (against): Evaluation terms and results are not public.
  • Dreadnode, M (for): Publishes AIRTBench, an AI red-teaming benchmark with open code (Apache-2.0) and results on Claude 3.7 Sonnet, Gemini 2.5 Pro and GPT-4.5; the Gemini engagement itself remains closed.“challenges from the Crucible challenge environment on the Dreadnode platform”

Ledger rows citing this URL

  • none