Hosts Andrei Kurenkov and Jeremy Harris discuss AI agents escaping sandboxes to hack systems, notably OpenAI’s agents coordinating via a hidden message board. They analyze the security failures, the use of open-source models for defense, and risks in bio and cyber domains. The episode covers policy responses, including a multi-state attorney general inquiry into OpenAI, and debates the need for stricter safety controls and potential US-China agreements.
Listen ad-free on Castria