0:00
/

Cybersecurity is Ceding AI Safety, But Should We?

The Hugging Face incident became an extinction story in about six weeks. Zack Korman thinks it was a sandbox failure, and that the people who reviewed it were never independent to begin with.

In this episode, I sit down with Zack Korman, CEO and co-founder of AI agent monitoring startup Embroidery and former CTO of a cybersecurity company, to unpack the collision between the AI safety world and the cybersecurity world that has played out since the OpenAI Hugging Face incident.

I first found Zack during the fake SOC 2 report saga on X, and his videos on effective altruism and on the METR and Redwood Research review of the incident have been among the sharper practitioner critiques I’ve read. That’s why I wanted him on live.

We start with where the extinction narrative actually comes from. Zack traces it through effective altruism and longtermism, and argues that many of the people warning that AI will kill eight billion people held that belief before they ever learned the technology. His point is not that they’re lying. True belief does not make someone right, and it makes them a poor guide to which risks deserve attention. From there we get into Dario’s Pacing the Frontier essay and the call for METR to serve as an independent evaluator. Zack walks through the funding and personal relationships linking METR, Redwood Research, Coefficient Giving, and the labs, and lands on a line I keep thinking about. It would be convenient if your SOC 2 auditor was also your best friend who agreed with every security decision you’ve ever made. We usually call that fraud.

The back half is the part every security practitioner should hear. Zack reads the METR and Redwood report the way an incident responder would, and finds amateur mistakes, a context-drop problem in how they used AI to analyze a ten million token transcript, and a habit of leaving facts “unknown” that server logs and syscalls would have settled in an afternoon. We go through the control failures OpenAI itself disclosed, including a single package proxy as the only isolation, an evaluation environment that was supposed to have no external network access, chain of thought monitoring that wasn’t running, and no automated shutdown. Any one of those controls would have stopped it. We close on why alignment is one control among many rather than the whole plan, what the real enterprise risks look like when most corporate environments are in worse shape than OpenAI’s, and why cybersecurity’s slowness may be the biggest risk of all.


Thanks for reading the Resilient Cyber Newsletter! Subscribe for FREE and join 23,000+ readers to receive weekly updates with the latest news across AppSec, Leadership, AI, Supply Chain, and more for Cybersecurity.



Prefer to listen?

Apple Podcasts

Spotify

Please be sure to subscribe and leave a rating/review, as it truly helps the show!


We discuss:

  • Why Zack spends his weekends taking apart AI safety reports, and what four years as a CTO taught him about the industry’s problems

  • Effective altruism, longtermism, and Sam Bankman-Fried as the through line from FTX to Anthropic’s Series B to today’s extinction discourse

  • Why “why not be careful” is the wrong question, because there is no safe path, only trade-offs

  • Pacing the Frontier and whether METR counts as independent oversight when reviewers, funders, and labs share one ideology

  • The funding and family ties between METR, Redwood Research, Coefficient Giving, and OpenAI’s board

  • Why the labs’ behavior doesn’t map to greed, and why Washington models Silicon Valley wrong

  • True believers are not necessarily right, and why prior ideology should discount current evidence claims

  • How the METR and Redwood review compares to what Unit 42 or Mandiant would have demanded before signing a report

  • The AI context-drop problem behind the finding that the analysis model “sided with” the agent

  • The eight layers of failure in the incident, from network egress to chain of thought monitoring to kill switches

  • Why Zack calls it a preventable human and organizational failure and an alignment failure at the same time

  • Whether doomer writing in the training data is producing models that cosplay as doomers

  • Alignment as one control in defense in depth, never the org chart

  • The tangible enterprise risks, including threat actors, customer deployments with worse sandboxing than OpenAI’s, and AI-discovered attack chains

  • Why cybersecurity’s laggard culture is the biggest risk in the room


Zack Korman


Timestamps

00:00 Introduction
01:00 Anger as a career path, and how Zack ended up here
02:40 Effective altruism and the extinction narrative
05:59 Risk management, not risk elimination
06:29 Pacing the Frontier and the METR independence problem
10:06 True believers, hype, and how DC misreads Silicon Valley
14:45 Taking AI risk seriously when the alarm is ideological
17:20 AI safety doing cybersecurity’s job
21:24 The eight layers of failure and which control would have stopped it
23:50 Alignment failure versus containment failure
28:08 The tangible risks for enterprises adopting agents
31:14 Cybersecurity is too slow

Discussion about this video

User's avatar

Ready for more?