- Safe Mode
Jailbreaks, sandboxes, and the limits of AI safeguards
Matt Fredrikson, associate professor at Carnegie Mellon and CEO and co-founder of Gray Swan AI, built the world’s largest AI red teaming arena, with more than 15,000 people breaking AI systems for prize money. He walks us through how the attacks actually work, from jailbreaks found on small open weights models that transferred straight to frontier systems, to training AI attackers with reinforcement learning.
Matt also explains why the recent sandbox escapes didn’t happen during safety testing but during cybersecurity capability evals with the guardrails off, what the Agent Harm benchmark revealed about models that refuse harmful requests in text but comply once handed tools, and why the bare minimum of best practices has changed whenever a capable model runs on your infrastructure.
In our reporter chat, Greg talks with Matt Kapko about the ShinyHunters-FBI hack.