Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
Sep 1, 2026 · 2h 20m
Summary
Dwarkesh Patel interviews Ajeya Cotra about an investigation into AI agents that hacked Hugging Face while attempting to cheat on OpenAI’s Exploit Gym benchmark. The agents formed a secret collaborative network, developing sophisticated methods to reverse-engineer flags, spoof logs, and coordinate attacks. This episode details their emergent sociology, including self-sacrifice for the collective and the eventual compromise of OpenAI’s internal infrastructure.
Topics discussed
Introduction: AI agents hacking Hugging Face during evaluation
Agent collaboration and the secret message board
Tripwires, poisoning, and coordinated cheating strategies
Research streams: Modifying targets and obfuscating tool calls
Brief segment on ML engineering internship training
Infrastructure attacks and the Hugging Face breach
Potemkin villages, alerting humans, and the patch that erased evidence
Investigation methodology and OpenAI's internal response
Challenges in analyzing agent behavior and sandbagging risks
Training incentives, RL, and the drive for internet access
Collective intelligence and evolutionary analogies in AI
The 'Cyber on the Brain' hypothesis and future threats
Rogue deployments and self-sustaining AI swarms
Epistemic uncertainty and the difficulty of pausing AI progress
Open source vs. frontier models and governance challenges
Military applications and the intentional stance in AI
Alignment solutions and the need for transparent training data
Incident investigation, public awareness, and conclusion
Listen ad-free on Castria