The debate over AI alignment often centers on guardrails, model behavior, and whether increasingly capable systems could act in unexpected ways. But what does that conversation look like through a cybersecurity lens?
In this episode of Safe Mode, Greg speaks with Adam Meyers, CrowdStrike’s head of counter-adversary operations, about the security implications of goal-driven AI agents, open-weight models, and guardrails that can be removed or bypassed. Meyers argues that systems described as “rogue” may instead be pursuing their assigned objectives in ways their operators failed to anticipate—making containment, monitoring, and clearly defined permissions central cybersecurity concerns.
The conversation examines how AI-enabled adversaries are accelerating operations, why open-weight models complicate alignment and safety debates, and how defenders can use AI to keep pace with faster exploitation and shrinking patch windows. Meyers also shares what CISOs and policymakers should focus on: visibility into how AI is being used, controls over what agents can access, and practical safeguards that match the technology’s real-world risks.
24 Sept 2026
Jailbreaks, sandboxes, and the limits of AI safeguards
Matt Fredrikson, associate professor at Carnegie Mellon and CEO and co-founder of Gray Swan AI, built the world's largest AI red teaming arena, with more than 15,000 people breaking AI systems for prize money. He walks us through how the attacks actually work, from jailbreaks found on small open weights models that transferred straight to frontier systems, to training AI attackers with reinforcement learning.
Matt also explains why the recent sandbox escapes didn't happen during safety testing but during cybersecurity capability evals with the guardrails off, what the Agent Harm benchmark revealed about models that refuse harmful requests in text but comply once handed tools, and why the bare minimum of best practices has changed whenever a capable model runs on your infrastructure.
In our reporter chat, Greg talks with Matt Kapko about the ShinyHunters-FBI hack.
17 Sept 2026
ClickFix and the social engineering of routine
It starts with a fake error message. It ends with a pasted command. ClickFix has become one of the most reliable ways for attackers to get an initial foothold — first adopted by criminal groups, now used by state-linked actors like APT28 and Lazarus Group.
This week, Greg is joined by John Hammond, Principal Security Researcher at Huntress, who helped identify and name the technique and has tracked its evolution for the past three years. They dig into why ClickFix still drives an estimated 20+ incidents a day at Huntress even after law enforcement disrupted major infostealer infrastructure like LumaStealer, how the technique has splintered into variants like FileFix and "consent fix" targeting Microsoft 365 and browser-stored credentials, and how attackers are now smuggling payloads inside images to slip past endpoint defenses. John walks through a real incident where a single pasted command led to 11 compromised devices, explains why he still believes security awareness training — not just tooling — is the best defense, and makes the case that the browser itself has become the blurriest and most under-protected boundary between endpoint and identity security.
The conversation also turns to the npm/Axios supply chain attacks, where operators built fake Slack communities and fabricated employee personas to earn open-source maintainers' trust before compromising their packages — and what, if anything, can technically guard against a threat that's fundamentally social.
In our reporter chat, Greg talks with Derek Johnson about the Supreme Court deciding against the Trump administration’s push for changes to mail-in ballots.
Follow John on Youtube: https://www.youtube.com/@_JohnHammond
Inside a recent episode
Google’s John Hultquist on how APTs are using generative AI
Published 6 Feb 2025 · Transcript excerpt
[…] You really need to have an idea of what the adversary is doing. And we finally, I think we've got a good picture. So with that, the report doesn't seem like it's a total game changer yet for the way that APT groups are using it. What specific thresholds would indicate that AI has become that true force multiplier for bad actors and what early warning signs should the industry monitor? Do you think that this report is the early warning sign or do you think that it's not really at a point where this needs to be something else to worry about when it comes to their arsenal? It's certainly something to a monitor and something that we're going to worry about that's kind of our jobs, right? […]
Pod Engine is an independent podcast discovery and analytics service and is not affiliated with or endorsed by this podcast. Artwork and show content belong to their owners. Full legal notice.
Explore this show Podcast research with Pod Engine