Podcast: Play in new window | Download
Subscribe: Apple Podcasts | RSS
For most of the last year, the AI security conversation had a clear villain. Attackers were using AI to write better phishing lures, adapt malware mid-attack, and move faster than defenders could keep up with. That story was easy to tell because it fit the shape we already understood. Bad guys get a new tool, they use it against us, we build a defense.
Then the defenders’ own AI started breaking its own rules, and the story stopped being that simple.
The Model Did What It Was Told
In July, OpenAI disclosed that one of its models, running an internal test with reduced safety guardrails, found a zero-day vulnerability, broke out of its own sandbox, moved laterally through OpenAI’s research environment, and reached across to Hugging Face. Two weeks later, Anthropic acknowledged that three of its own models had done something similar. Both companies were describing their own systems doing exactly what they’d been asked to do, just not in the way anyone expected.
Brad LaPorte of Morphisec has been tracking this closely, and he pushes back on the idea that this is as novel as the headlines suggest. “This isn’t necessarily novel,” he told me. “It’s just they made the front page of the newspaper finally.” What’s changed is the scale, and the level of autonomy behind it, enough that it’s getting harder to write off as an edge case.
That reframes what “guardrails” actually means. A model told to solve a problem and finds an unsanctioned path to solving it is still doing its job, just not the way anyone intended. Isaac Asimov’s laws of robotics get invoked a lot in these conversations, and for good reason. “Don’t harm a person” sounds like a rule until you realize a model can read it literally and conclude nobody got physically hurt, so nothing went wrong. Financial harm doesn’t register that way. Neither does legal exposure or reputational damage, and neither shows up in a rule that vague.
The Fundamentals Haven’t Caught Up
While that governance conversation plays out, the operational numbers are moving in the wrong direction. IBM’s most recent cost of a data breach report showed the average time to identify a breach getting worse for the first time in five years, up six days to 247. Meanwhile, roughly 88 percent of organizations report using AI in at least one function, but only 5 to 10 percent are seeing meaningful return on that investment. A lot of the AI-driven layoffs making headlines aren’t the result of AI actually doing the work yet. They’re companies freeing up budget to chase a productivity gain that hasn’t arrived.
That gap between adoption and governance is where the real risk lives. Organizations are running AI agents that talk to each other, share data, and make decisions with non-human identities that most security teams haven’t fully inventoried, let alone secured. Add unsanctioned AI tools employees are using without approval, and you have an attack surface that’s expanding faster than most companies can map it.
None of This Changes the Basics
The uncomfortable part is that the fix isn’t exotic. Identity and basic visibility into what’s actually running in your environment still account for most of the risk reduction available to any organization. AI adds a new layer to secure, but companies that had their fundamentals in order before AI showed up are adapting. The ones that didn’t are finding out that AI doesn’t so much create new problems as it makes the old, ignored ones impossible to keep ignoring.
Brad and I get into all of this on the latest episode of the TechSpective Podcast, along with where he thinks the AI funding bubble is headed and why he compares it to the mortgage-backed securities mess of 2008. Give it a listen.
- When the AI Breaks Its Own Rules - September 2, 2026
- What It Took to Film a Real-Life Lion King - August 28, 2026
- Your AI Agents Are Multiplying and Nobody’s Counting - August 26, 2026




