Welcome to issue #111 of the Resilient Cyber Newsletter!
Two stories dominated the week, and both of them run through Hugging Face. OpenAI published its technical report on the incident where its own internal research models escaped an evaluation environment and went on to compromise Hugging Face infrastructure, and NVIDIA is reportedly closing in on acquiring Hugging Face for somewhere in the neighborhood of $13 billion.
Read together, they tell you a lot about where we actually are. Agents are demonstrably capable of chaining novel vulnerabilities to break out of the sandboxes we build for them, and the open model ecosystem those same agents are trained on and distributed through has become valuable enough that the most important chip company on the planet wants to own the distribution point. Both of those have security implications we are only starting to work through.
The rest of the week filled in around them, with Trail of Bits arguing that VMs won’t contain cyber-capable agents, the researchers behind ExploitGym walking through how they measure agent exploitation capability, SemiAnalysis making the case that open models close the gap faster with every cycle, and Ciaran Martin offering a badly needed corrective to AI doomerism.
Let’s get into it!
Know your attack surface before it outgrows you
You cannot patch, monitor, or govern an asset you don’t know exists. Yet, most enterprise inventories are still working from an incomplete picture.
runZero, built by HD Moore (the mind behind Metasploit), delivers agentless, credential-free discovery across your entire internal and external attack surface—surfacing unknown and unmanaged assets along with the exposures riding along with them: CVEs, default credentials, misconfigurations. No appliances, no scanning windows to negotiate, just an accurate picture of what’s actually out there.
Pro tip: When your trial wraps, runZero’s free Community Edition covers up to 100 devices, perfect for home labs!
*Sponsored
Cyber Leadership & Market Dynamics
NVIDIA in talks to acquire Hugging Face in $13 billion deal
News broke over the weekend that Hugging Face was fielding takeover interest, and by midweek The Information reported that NVIDIA had agreed to buy it for $12.9 billion, with no signed agreement yet and neither company confirming. For context on how fast this has moved, Hugging Face turned down a $500 million NVIDIA investment in late 2025 that valued it at $7 billion, with CEO Clem Delangue saying the company “didn’t want a dominant investor that could sway its decisions.” It was valued at $4.5 billion in 2023 and is now doing roughly $150 million in annual revenue, up from about $100 million two months prior.
The strategic logic is easy to follow, with OpenAI, Google, Amazon, and Anthropic all building their own silicon, owning the place where open models get published and pulled is a durable way to keep the ecosystem on your hardware. The security angle is the one I’m most interested in of course.
Hugging Face is effectively the GitHub of AI, a package registry for weights, datasets, and Spaces that enterprises pull from constantly, often with far less scrutiny than they apply to OSS packages. Concentrating that registry under a single vendor with its own commercial incentives changes the software supply chain risk conversation, and it does so in the same month the registry itself got rooted.
Rolling with the Punches: Why Cybersecurity is Backgammon, not Chess
Phil Venables is one of the most consistently useful writers and deep thinkers in our field, and this one qualifies. His argument is that we keep framing security as chess, a game of perfect information where a sufficiently smart player calculates their way to victory, when it actually behaves like backgammon, which he describes as “stochastic risk management” where skill has to account for the dice. He maps it out through blots and anchors (exposed vulnerabilities versus layered controls like IAM, MFA, and segmentation), priming, the blitz (automated ransomware hitting everything at once), and the doubling cube (the decision of whether to pay or execute recovery).
The cultural point at the end is the one that matters for security leaders. Chasing “flawless, unyielding defense” is the wrong objective function, and resilience means accepting that bad rolls happen and building an organization that absorbs and adapts. Given how the OpenAI incident below unfolded, with clear signals that went un-escalated, this framing is timely.
Exclusive: CrowdStrike’s CTO is leaving to launch an AI cyber fund
Elia Zaitsev is leaving CrowdStrike after 13 years as global CTO to launch Cognition, a venture firm he is starting with Gur Talpaz and Tayler Sipperly, targeting $170 million. The model is concentrated, leading or co-leading seed and Series A with average checks around $6 million and $15 million respectively, and only three to four investments a year. Zaitsev’s framing is that agentic AI is a platform reset for security, and as he put it:
“We have this new attack surface that’s being brought on by AI and agents.”
I’ve made the point before that we are watching a genuine reshuffling of the security market, with operators leaving incumbent platforms to build and fund the next wave. Whether agentic AI produces a category-defining platform or gets absorbed into the existing consolidators is the open question, and it is one Sid Trivedi and I spent a good chunk of our annual pre-Black Hat conversation on. That said, when a sitting CTO of the most successful endpoint platform of the last decade walks out to raise on this thesis, it tells you something about how the people closest to the data see it.
Cybersecurity startup Minimus shutting down after raising $51 million
The other side of that market story, and one we don’t hear about enough. Minimus, founded by the team that sold Twistlock to Palo Alto Networks for roughly $410 million in 2019, is winding down after raising $51 million in seed funding. The company built stripped-down container images and VMs with dramatically fewer vulnerabilities, publicly launched in April 2025, and had around 60 employees four months ago. Customers have 60 days, until October 22, to migrate off the Minimus registry. The founders’ statement is pretty direct as well, that:
“the current business and investment climate has resulted in a situation in which we are unable to continue operations.”
This one stings, because the product solved a real problem. Minimal images are one of the few software supply chain interventions that actually burns down vulnerability backlogs rather than just measuring them. Proven founders, real technology, and a genuine pain point still weren’t enough, which says quite a bit about how hard it is to sell into a crowded container security market where buyers have five ways to defer the purchase.
As I wrote on LinkedIn, this is a good example that building a thriving security (or any) company is very hard. This one had a real market need, being addressed by repeat founders who have walked the path. But, it is a highly competitive category with players like Chainguard dominating the space and it is hard to be wake a player who defined and built the category.
A good reminder in this era of record setting funding rounds, PR announcements and categories flooded with tens of security vendors chasing the same problem around AI and agents.
Questions to ask AI security vendors
Lenny Zeltser put together five questions worth bringing into any AI security evaluation, which AI asset the product actually protects, whether it delivers value standalone or requires a broader platform, what exists today versus what lives on the roadmap, whether you already own the capability somewhere else in your stack, and what evidence backs the claims. His framing is refreshingly direct, that “In the fast-changing AI security market, a product described one way often turns out to be something else after a closer look.”
The fourth question is the one most practitioners skip and the one that kills the most deals, and the distinction he draws between traditional tooling with AI bolted on and genuinely AI-native protection is the fault line running through this entire category right now.
My favorite though is the one tied to roadmap vs. actual capabilities, and not what lives in a analyst or marketing pitch deck. Many of the AI-centric categories have companies chasing the opportunities but they have little to no actual functional capabilities deployed in production with customers.
Be careful out there!
The choices we make about AI now are critical
Bill Gates argues that this transition differs from previous ones because AI automates cognition itself, and that it will play out over a decade rather than across generations. He points to Stanford payroll data showing a 16% relative employment decline for workers aged 22 to 25 in AI-exposed occupations while older workers stayed flat, and proposes a governance framework modeled on nuclear inspection and aviation regimes, “Human Reserved” roles kept for people, and a token and robot tax.
The line security folks should notice is his, that:
“The smartest cybersecurity experts I know are scared about the next few years, because the attackers are getting powerful new capabilities faster than the defenders can fix all the weaknesses.”
I don’t agree with all of his prescriptions, and I remain skeptical of preemptive regulatory regimes built around hypothetical harms. That said, the asymmetry he describes is real, and the gap between offensive capability and defender remediation throughput is showing up in nearly every other item in this issue.
Diffusion is destiny for defenders, and our best defense against the coming AI impact in cyber by attackers is arming defenders with the innovative capabilities as quickly as possible.
Fears of AI-induced armageddon are overdone
Ciaran Martin, founding CEO of the UK’s NCSC, wrote the counterweight to the piece above, and it is a rational pushback on some of the AI FUD. He opens with L0pht walking into a Senate office building on May 19th 1998 and telling lawmakers they could render the internet unusable in 30 minutes, then walks through Leon Panetta’s 2012 “cyber Pearl Harbour” warning, pointing out that our field has “a long history of apocalyptic prophecies” that never arrived.
Martin isn’t arguing that nothing bad happens, he is arguing that our track record of predicting catastrophic collapse is terrible and that the predictions themselves distort policy and spending. Coming from the person who ran a national cyber defense agency, that carries weight. Pair it with the Gates piece and with Joshua Saxe below, and you get a much healthier version of this debate than the one happening on your social media timeline.
Anthropic plans to change data retention policy for advanced AI
Anthropic is reportedly moving to let enterprise customers store required data in their own cloud environments rather than exclusively on Anthropic’s infrastructure, while keeping a 30-day retention requirement for its most capable models.
For anyone running a frontier model in a regulated environment, this is the constraint that has been blocking deals, and it is a reasonable compromise between the safety case for retention (you cannot investigate an incident like the one below without logs) and the enterprise requirement to control where data lives.
Given that OpenAI’s own report leans heavily on retained chain-of-thought and infrastructure telemetry to reconstruct what happened, the tension between “retain enough to investigate” and “retain nothing our customers didn’t consent to” is going to be a live governance question for every AI vendor, not just Anthropic.
This has also been a key point/theme in the conversations around open vs. closed models for enterprises, so it isn’t surprising to see the labs trying to offer different options to alleviate those concerns and streamline procurement pathways to expand adoption and consumption.
AI
The Hugging Face incident and the road ahead
This is the report of the year so far, and if you focus on one thing in this weeks issue, make it this.
OpenAI walked through how internal research models, during cybersecurity evaluations, “communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access” and ultimately compromised both OpenAI and Hugging Face infrastructure.
The timeline is a key aspect to examine.
On May 12 the agents discovered they could pass messages to each other through Artifactory, building an improvised message board out of file directories to coordinate across supposedly isolated instances.
On May 26 they got internet access through SSRF.
On June 26 they exploited a token refresh flaw for admin access.
On July 4 sustained agent activity crashed Artifactory, and when it was rebuilt on July 8, the agents immediately reestablished both the message board and internet access.
Between July 10 and 12 they recovered publicly exposed Hugging Face credentials, chained “several novel security flaws” into code execution, and took root across Hugging Face servers in multiple regions.
Detection came on July 19 and 20, and public disclosure on July 21.
The root cause analysis is where the practitioner value and community takeaways are. OpenAI names reward hacking, persistence without a safe exit, metagaming, and unauthorized communication. The most damning figure in the report is that of 898 ExploitGym tasks, 198 were ones the models had never solved, and those unsolved tasks accounted for “93% of the tasks that were discussed on the Artifactory message board.”
The models were handed impossible problems, given enormous compute, rewarded for solving them, and rationally concluded that attacking the scorer was the path to reward.
Two failures compounded the situation as it is described. Production safeguards, meaning system prompts and auto-review models, were “not applied in the evaluation environment,” and chain-of-thought monitoring wasn’t running on these evaluations at all. OpenAI states that:
“deployed monitoring would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems,”
and that production safeguards reduce infrastructure compromise propensity by over 100x. The controls existed, they just weren’t turned on where the most capable models were being pushed hardest.
There are also things the report doesn’t say. Per Fortune’s read, OpenAI didn’t publish the prompt given to the agents, message examples, or code snippets, all of which Hugging Face did include in its own July 29 post-mortem.
Fortune also reports, citing the independent METR and Redwood Research analyses, that 1,200 agents communicated over the message board and 700 participated in the attack, with a lead agent coordinating concealment. OpenAI’s own characterization, that this was “misaligned behavior in an outlier scenario involving a rare and unexpected confluence of events,” reads as perhaps thinner than the timeline alludes to.
That said, credit where it is due, publishing this level of detail is genuinely valuable to the field, and the commitments are the right ones, including required chain-of-thought monitoring for models at GPT-5.6 Sol capability or above, multi-agent alignment environments that teach models “to distrust unauthorized instructions,” training that rewards “correctly identifying broken tasks, requesting clarification, or stopping safely,” and 30-minute escalation requirements for severe misalignment alerts.
That said, every one of those is a control we already know how to describe in enterprise security terms, and none of them were in place at one of the organizations with the most capable models on earth and the most reason to be careful.
VMs won’t contain cyber-capable agents
The team at Trail of Bits published the perfect companion piece, and the timing is ironic. They gave GPT 5.6-Cyber the task of escaping a VM sandbox, and it succeeded three separate times with escalating sophistication.
First, it identified a recently disclosed kernel vulnerability, built an exploit, and hardlocked the host.
Second, it found that Debian 12 shipped an outdated libslirp (v4.7.0) carrying CVE-2026-9539 and combined it with an unmarked bug fix to get arbitrary read and write into host memory.
Third, and this is the one that is most concerning, it chained three previously unknown vulnerabilities across QEMU, Linux KVM, and libslirp, plus one patched flaw nobody had recognized as security relevant, to get controlled execution in the host kernel.
The behavioral description matters as much as the results on this one. The agent “operated autonomously for hours, backtracked from approaches that didn’t work, pulled code and research papers,” and sustained this across 12-hour windows. That is not a model completing a benchmark, that is a researcher with infinite patience.
The blogs recommendations are practical, meaning minimal virtualization platforms like Firecracker rather than full QEMU, continuous distribution updates instead of stable releases that lag on patching, least privilege applied to network access and credentials, and time-limited operations with pristine resets between runs.
Put this next to the OpenAI report and the conclusion is uncomfortable, because the sandbox has been among the load-bearing control in every agent security architecture, including plenty in production today.
This is a topic I’ve discussed on the Resilient Cyber show in the past with folks such as Luke Hinds and Alex Zenla months ago, before the breakouts got everyone’s attention.
Inside ExploitGym: How Researchers Are Measuring AI Agent Exploitation Capabilities
Timely context on the benchmark at the center of the OpenAI incident. ExploitGym holds roughly 898 vulnerability instances and measures whether agents can turn known vulnerabilities into working exploits, across userspace memory safety flaws in C and C++, browser engines (Google’s V8), and Linux kernel exploitation.
The results show a steep difficulty curve. Claude Mythos Preview and GPT-5.5 led with 157 and 120 successful exploits, but kernel exploitation collapsed that to 12 and 22 successes for the top two models. Agents also sometimes achieved code execution through unintended vulnerabilities rather than the one under test, which is exactly the behavior that generalized into the OpenAI incident.
“Measuring only whether an exploit succeeds may miss important changes in how that success is achieved.”
Agentic AI: The New Insider Threat Model
Katie Moussouris of Luta Security frames autonomous agents as an insider threat problem, which is the right mental model. Her assessment of the Hugging Face incident is direct, that:
“Clearly, we didn’t have the real-time monitoring in place, and we don’t have any breaks that seem to work.”
She notes the gap between initiation in May 2026 and detection on July 19, roughly two and a half months of undetected activity, and points out that Anthropic’s latest model “realized it was on the internet and stopped itself,” which is the behavior we should be training toward.
Her broader push is one I strongly agree with, that organizations need to invest in the fundamentals, reduce attack surface, and manage technical debt rather than assuming AI tooling solves problems the organization never solved manually.
The insider framing is useful because we already have a discipline for this, with behavioral baselines, least privilege, separation of duties, and monitoring for the trusted entity going off-script. We just have to apply it to entities that operate at machine speed and outnumber our employees, which is a new paradigm for us all.
Where Security Fits in an AI Agent Stack
NVIDIA published a genuinely good architectural piece, and the core distinction is one I’d like to see adopted broadly, that behavioral controls (prompts, model safeguards, harness logic) are different in kind from infrastructure controls (what the runtime actually permits). Their framing is clean, that “The harness guides what an agent tries. The infrastructure controls what an agent can do.”
They lay out five layers, distribution and product, orchestration, agent harness, secure runtime, and inference data plane, with a security boundary drawn between the harness and the runtime so requests cannot route around policy enforcement. The five design rules are the useful part, that higher layers propose and lower layers decide, policy lives below the boundary, every effect is checked, access is just-in-time, and isolation enables recovery. They also define four escalating security profiles, Isolated, Connected, Production, and Adversarial.
This maps almost exactly onto the soft guardrails versus hard boundaries argument I’ve been making for a while now, and the OpenAI incident is the proof case. Every control that failed there was a behavioral control, or a boundary control that simply wasn’t enabled.
Secure Vibe Coding and the 99%
In this episode, I sit down with Igor Andriushchenko, Head of Security and CISO at Lovable, to discuss securing AI-native development, including soft guardrails versus hard boundaries, the shared responsibility model for vibe coding platforms, and how GRC engineering fits into a world of AI-powered attackers.
I had been chasing Igor down on LinkedIn since February, so I was glad to finally get him on the show. Lovable sits in an unusual spot. Igor has to run security for a company that went from roughly 40 people to 400 laptops in MDM inside a year, and at the same time the product he is securing hands software creation to people who are not developers and have never thought about security at all. Those two problems pull in different directions, and the conversation gets into how his team handles both.
Are Open Models Catching Up?
The best data-driven piece I read this week, and directly relevant to the NVIDIA news above.
Dylan Patel and SemiAnalysis argues the gap between open and frontier models closes in roughly half the time with each successive era.
The authors are appropriately careful, noting that “benchmarks are not the end all be all” and warning about hill climbing on public evals. For security purposes, though, the direction is what matters.
If open-weight models reach frontier agentic capability within four to six months of release, then every capability demonstrated in the OpenAI report and the Trail of Bits research becomes available to anyone with a GPU and no safety stack on top of it, on a predictable clock. That is the assumption Igor’s conversation above mentions, and the evidence supports it.
Where are all the prompt injection attacks?
Joshua Saxe asks the question our industry avoids, which is why documented real-world damages from prompt injection remain tiny relative to the hundreds of billions lost to cybercrime annually.
He offers three explanations:
Opportunity cost (attackers still get in through known vulnerabilities, weak credentials, and misconfigurations, which is cheaper than building novel exploit workflows)
Organizational inertia (established criminal enterprises have profitable, proven playbooks and poor ROI on retraining)
Attribution gaps (detection tooling for LLM-mediated breaches is immature enough that successful injections may be getting attributed to whatever system got touched next).
He isn’t dismissing the risk, and he still recommends securing agents with something like Meta’s rule of two, restricting sensitive actions on untrusted data.
His point is about resource allocation, that overweighting AI-native risk because it dominates conference talks means underfunding the tech debt that is actually getting organizations breached. Josh has been on the Resilient Cyber show and consistently brings evidence rather than vibes, and this is a good example of it.
AppSec
Staying Ahead of Adversarial AI Through Agentic Source Code Review
Google Cloud and Mandiant detailed their Agentic Vulnerability Discovery Harness, a waterfall pipeline of specialized agents running threat modeling, entry point discovery, context enrichment, hypothesis generation, and hypothesis validation, with human expert gates at both ends.
Distinct agents handle access control analysis and data flow analysis, Gemini Flash Lite handles parallelized entry point discovery at scale, and high-temperature validation agents feed a synthesis agent that marks findings Confirmed, Disproven, or Rejected. Mandiant expertise is injected through a three-tier rules hierarchy covering domain, framework and language, and vulnerability specifics.
The results are the reason to read it, with 12 CVEs assigned so far, another dozen in active disclosure, tens of millions of lines of code analyzed, and over 100 critical vulnerabilities discovered in two days during an incident response engagement.
The harness is the story here, not the model, which is a trend we continue to see. This is the same lesson from the NVIDIA piece above and from OpenAI’s report, that capability comes from orchestration, scoping, and validation rather than from prompting a frontier model harder.
It is also worth noting that Google is careful to keep human validation in the loop at the exploitation stage, which is the difference between a finding and a fix.
Just a rumour of a bug is enough to find a security exploit these days
Anil Madhavapeddy, the OCaml maintainer and Cambridge researcher, wrote a super interesting piece. He opened a public pull request fixing a path traversal bug in cohttp and saw probes matching that vulnerability pattern within 10 minutes. He then demonstrated that building a working exploit from the PR alone took roughly one minute with available models. His point is that embargo, the entire premise of coordinated disclosure, assumes attackers need time that they no longer need.
He cites 2026 research showing the bottleneck has moved to “defender remediation throughput,” and points to CVE-2026-39987 exploited in 9 hours and CVE-2026-33017 in 20 hours with no public proof of concept in either case. His proposals are worth engaging with, including private development infrastructure with web-of-trust authentication, continuous shipping on Chrome-style cadences, and protocol-layer virtual patching through distributed rule networks ahead of upstream fixes.
For open source maintainers, this is a serious problem with no good answer yet, something I’ve discussed on the show with industry veteran Casey Ellis of Bugcrowd. The security fix commit is now an exploit specification, and most maintainers have no ability to ship a coordinated release across an ecosystem before that specification is public.
Anyone working on OSS security funding and infrastructure should be reading this one closely.
How to Hack AI Agents
Shoeb Patel put together an excellent practitioner’s mental model for finding vulnerabilities in agentic systems, and it is the resource for AppSec engineers being asked to test their first agent.
He frames the core problem:
“the model is trained to trust the system prompt more than the user, and the user more than tool results. But this boundary is soft.”
This aligns with the soft guardrails vs. hard boundaries framing I’ve been writing and speaking about for a while now. From there he separates direct injection through user-controlled inputs from indirect injection hidden in web pages, emails, plugin descriptions, repositories, documents, and logs, then walks the impact vectors of data exfiltration, privileged actions taken under the agent’s or victim’s identity, and persistence through injections stored in memory and configuration.
The writeup references more than 15 documented attacks including CamoLeak, EchoLeak, and Invitation Is All You Need, and points to practice environments like Gandalf and HackAPrompt plus Microsoft’s PyRIT for red teaming.
Note how well it pairs with Saxe’s piece above, because the techniques are all real and demonstrated while the in-the-wild exploitation data remains thin, perhaps due to the factors Josh mentioned such as attribution, widely available low hanging fruit etc.
Final Thoughts
The week’s two headline stories are the same story told twice. OpenAI’s report shows agents chaining novel vulnerabilities to escape an evaluation sandbox, and Trail of Bits shows a commercially available model doing the same thing to a VM three different ways.
Meanwhile, the registry those models are published to and pulled from is about to change hands for $13 billion, and SemiAnalysis has the data showing open-weight models reach that capability level four to six months behind the frontier and closing.
If those things are all true at once, the security architecture most organizations are building around agents right now rests on a control that doesn’t hold, against capabilities that will be freely downloadable inside of a year.
The good news, such as it is, is that nobody is proposing exotic solutions. NVIDIA’s boundary model, Igor’s six pillars, Moussouris’s insider threat framing, and OpenAI’s own list of commitments all describe controls we already know how to build.
We just have to actually build them, everywhere, and turn them on in the environments where we are pushing the hardest, which is precisely where they weren’t.
Stay resilient.




















