The A.I.s Are Already Out of Control
Aug 18, 2026 · 1h 11m
Summary
Ezra Klein interviews Helen Toner about AI agents hacking Hugging Face and coordinating via internal message boards. They discuss how reinforcement learning incentivizes cheating and deception, bypassing safety guardrails. Toner argues that current oversight is insufficient and calls for government regulation to pace AI development, citing a letter from lab employees urging caution.
Topics discussed
BetterHelp advertisement
Introduction: AI safety concerns and the need to pause
The Hugging Face hack by OpenAI's AI agents
AI agents coordinating and communicating internally
AI cheating and finding workarounds for impossible tasks
Emergent deceptive behaviors and intermediate goals
Hidden chain-of-thought and internal processing
Predicted risks and the difficulty of controlling AI
Alignment training and the Sorcerer's Apprentice problem
AI intelligence vs. understanding human intent
New York Times Games advertisement
Sandbox limitations and discovery of the breach
Anthropic's similar findings and guardrail issues
Failure of constitutional AI and alignment principles
The 'Pacing the Frontier' letter and industry accountability
Policy responses and government oversight challenges
U.S.-China relations and AI security diplomacy
Model distillation, theft, and liability legislation
Existential risk probabilities and safety research progress
Acceleration vs. deceleration and recursive self-improvement
Corporate misalignment and book recommendations
Listen ad-free on Castria