0:00
/

Why Restricting AI Makes Us Less Secure

Frontier model restriction, the jagged frontier, Chinese open weights, and why AI cybersecurity will be won through defender adoption.

In this episode, I sit down with longtime AI and security leader Joshua Saxe to discuss why restricting frontier AI in the name of safety actually makes us less secure. Josh spent 15 years applying machine learning to security, built and ran the ML program at Sophos, and most recently led security for Llama at Meta before leaving to co-found a startup reimagining vulnerability and exposure management with agents.

I first heard Josh speak at Unprompted earlier this year and have been following his Substack ever since. He’s been writing some of the most cited pieces on the collision of AI, cybersecurity, and national security policy, and this conversation digs into the core of his diffuse or lose argument.

We chatted about:

  • Josh’s path from teenage blackhat to high school teacher, defense and intelligence work, Sophos, Meta, and now his own startup

  • Why restricting access to the best American frontier models harms defenders more than attackers

  • How monitored closed models put threat actors at a disadvantage, and why pushing them to self-hosted open weights blinds defenders

  • The jagged frontier, and why bug finding is a small slice of what attackers actually use AI for

  • The national security and supply chain risks of the world running on Chinese open weights models

  • Why exploits don’t cause cyberattacks, and the gap between CVE volume and actual exploitation

  • The dual use ceiling on classifiers and guardrails

  • Where defenders get the most from AI adoption right now, and using agents to burn down security technical debt



Prefer to listen? The episode is also available on:

Spotify and Apple Podcasts


Thanks for reading the Resilient Cyber Newsletter! Subscribe for FREE and join 20,000+ readers to receive weekly updates with the latest news across AppSec, Leadership, AI, Supply Chain, and more for Cybersecurity.


Attackers in the panopticon

Josh’s core argument starts with how the closed labs actually operate.

Every major lab runs inline guardrails plus detection and response teams monitoring traffic, and those are the same teams publishing the threat intel reports we’ve all read from Anthropic, Microsoft, and OpenAI.

That monitoring changes the calculus for any serious threat actor weighing a frontier model against an open weights alternative. As Josh put it, “You use Fable and some team in Anthropic might catch you in your tradecraft and then send that directly to the NSA.” A rational attacker organization stands up its own GLM 5.2 inference instead, fine-tunes away the guardrails, and trains on its own trajectories.

Restriction doesn’t take the capability away from attackers. It just moves their usage somewhere we can’t see it, while defenders lose access to the best tooling. That maps to a lesson enterprise IT learned the hard way with shadow IT, and we’re now repeating it at the national policy level.

Bug finding is maybe five percent of the story

Josh thinks the policy conversation is dramatically over-indexed on frontier models finding subtle bugs somewhat faster under specific harnesses and inference budgets. The killer app for attackers so far has been social engineering, which older and open models have handled for years, alongside target research and malware coding.

“The idea that we would block American companies and American leadership over this one capability, which is maybe five percent of what the models are useful for attackers, just seems very strange to me.”

I raised the Jagged Frontier work and research showing smaller and older models finding zero days, and Josh agreed the restriction argument falls apart once you look at what attackers actually do rather than what capability just shipped. His piece on exploits not causing cyberattacks makes the same case.

We’ve had Turing test passing chat models since 2022, and the predicted social engineering apocalypse never showed up at scale. I noted the FIRST mid-year data showing CVE volume climbing toward 70,000 while exploitation stays roughly flat. The bottleneck for most attacker constituencies was never finding bugs.

The dual use ceiling on guardrails

Given his time building safety classifiers at Meta, including the open sourced Purple Llama work, I asked Josh about the classifier-heavy approach we saw in the Fable redeployment. His answer was blunt.

Security is an inherently dual use domain, so there’s a very low theoretical ceiling on blocking attackers without also blocking the researchers and defenders doing legitimate work.

“If you’re gonna block attackers, you’re gonna block defenders with these guardrails.”

His own startup builds vulnerability discovery tooling and can’t fully use frontier models for exactly this reason. That’s the self-inflicted wound in a nutshell. The people most reliably stopped by these guardrails are the ones we most need to be helping.

Burning down the mountain of security tech debt

On where defenders should actually apply AI, Josh pointed to access management and over-permissioning, SOC automation, and above all the security technical debt every large organization carries.

Open admin portals without MFA, dangling dev servers, and known vulnerabilities that never got closed. Pre-AI, security orgs were maybe three percent of the company, reduced to nagging engineering and IT to fix what they found.

His startup is building agents that own discovery, triage, comms with issue owners, and verification of the fix, including agent-to-agent communication with the engineer’s own coding agent over MCP. Having spent much of my career in AppSec and vuln management, this resonated.

So much of the job has been relationship building and chasing people rather than technical security work, and agents may finally change that equation.

Reasons for optimism

Josh closed with the structural advantages defenders hold. AI that can surveil every log line an enterprise emits, the ability to find and fix bugs before software ever ships, and sheer numbers.

“There’s a thousand times more people interested in finding and fixing the bugs than there are in finding and exploiting them.”

If it’s just a question of tokens, the remaining problem is inter-organizational politics and community structure, not capability. I share that optimism, with the caveat that incentives, bureaucracy, and speed to market pressures will make the transition bumpy before the net positive shows up.

A big thanks to Josh for coming on.

Go follow his Substack for some of the sharpest writing on AI and cyber policy, and keep an eye on what he and his co-founders are building.

Discussion about this video

User's avatar

Ready for more?