Advertisement
  • Safe Mode

Jailbreaks, sandboxes, and the limits of AI safeguards

Matt Fredrikson, associate professor at Carnegie Mellon and CEO and co-founder of Gray Swan AI, built the world’s largest AI red teaming arena, with more than 15,000 people breaking AI systems for prize money. He walks us through how the attacks actually work, from jailbreaks found on small open weights models that transferred straight to frontier systems, to training AI attackers with reinforcement learning.

Matt also explains why the recent sandbox escapes didn’t happen during safety testing but during cybersecurity capability evals with the guardrails off, what the Agent Harm benchmark revealed about models that refuse harmful requests in text but comply once handed tools, and why the bare minimum of best practices has changed whenever a capable model runs on your infrastructure.

In our reporter chat, Greg talks with Matt Kapko about the ShinyHunters-FBI hack.

Weekly

Safe Mode

Every week we break down the most pressing issues in technology, provide you with the knowledge and tools to stay ahead of the latest threats and take you behind the scenes of the biggest stories in cyberspace.

Advertisement