Welcome to issue #110 of the Resilient Cyber Newsletter!
Every so often a piece comes along that gives you the frame for everything else you are reading, and this week that piece was Dan Lahav’s “The End-State Fallacy.”
His argument is that our industry keeps debating whether AI will ultimately favor offense or defense, and in doing so we skip past the part that actually determines what the next few years look like. As he puts it,
“even if the end state turns out to be defense-dominant, the next few years can still be sharply offense-dominant.”
The equilibrium is not the experience, the path is.
Once you have that frame in hand, this week’s stories stop reading like unrelated headlines and start reading like data points on a single curve.
Let’s get into it!
Cyber Leadership & Market Dynamics
Is the Cyber Industry Too Big?
Alan Shimel of Techstrong wrote this one from the floor of Black Hat USA 2026, and it is a great analysis of our space, and the numbers underneath it are worth your time.
Cyber generated roughly $335.8 billion in 2025, up from $214.9 billion in 2022, with projections of $371 billion in 2026 and $404.5 billion in 2027, but growth is moderating from 17.9% in 2023 to a projected 9% in 2027. There are 4,100+ vendors selling around 11,000 products, the top 10 vendors hold only 29.5% of the market, and Microsoft, the largest of them, holds under 10%. Black Hat itself drew 20,000+ attendees and 400+ solution providers at an average booth cost of roughly $250,000.
Alan’s answer to his own question is nuanced, and I largely agree with it. The industry is not too big for the problem, it may be too big for its own commercial plumbing. His observation that:
“the company with the best booth is not necessarily the company with the best product. It may simply be the company with the most capital to convert into visibility”
This is one every practitioner walking a show floor already knows in their gut. The Futurum survey he cites, with 929 decision makers, found 67% expecting budget increases while 42% plan to reduce vendors and 36% plan to add them, which tells you the consolidation story is far messier than the platform vendors would like, and per an IBM and Palo Alto Networks study he references, the average organization is managing 83 security solutions from 29 vendors.
His warning that “the answer to tool sprawl cannot be agent sprawl” is the one I would tattoo on the industry’s forehead heading into 2027. Read this next to the end-state fallacy piece further down, because 83 tools from 29 vendors and six month procurement cycles are exactly the institutional metabolism that makes the transition period offense-dominant regardless of how good our defensive technology gets.
Palo Alto’s CEO bought the dip. Five months later, his $10 million bet is worth $26 million
Nikesh Arora bought 68,085 Palo Alto Networks shares on March 27 at $146.88 per share, roughly $10 million, at a moment when the stock had fallen 20% since the start of the year on fears that AI would disrupt incumbent security vendors. That position is now worth over $26 million. PANW sits at roughly $315 billion in market value, closed the ~$25 billion CyberArk acquisition in February, and posted $3 billion in Q3 revenue at 31% year over year growth.
I include this less as a stock story and more as a signal about the disruption thesis. The market briefly priced incumbents as AI roadkill, and the incumbents responded by buying their way into the AI security stack, with five AI-related acquisitions in the past year alone. Arora’s own framing, that “as AI becomes more pervasive across the enterprise, it expands the attack surface area,” is exactly the pitch every platform vendor is now making. Pair this with the Shimel report above and you get the full picture of where the capital is flowing and why.
I recently came across a good long-form interview with Nikesh on the Sourcery podcast by Molly O’Shea
Anthropic’s annualized revenue surges to $65B
Per Marina Temkin at TechCrunch, Anthropic hit $65 billion in annualized revenue at the end of July, up from $47 billion in May and $9 billion at the end of 2025. That is $18 billion in annualized revenue added in two months. The company is projected to finish 2026 somewhere between $100 and $120 billion, against OpenAI’s $40 billion run rate, and both have filed confidential IPO paperwork.
To put into perspective what that means for our industry, the entire global cybersecurity market is around $371 billion projected for 2026 per the Techstrong figures above. A single model provider is on a trajectory to reach a third of that inside a year. This is the scale of the platform shift security is now being asked to secure, and it should inform how much of your program you are willing to bet on governance patterns that assume you get to move at your own pace.
Ramp AI Index, August 2026
Ramp’s Ara Kharazian titled this edition “Cracks in the AI Thesis,” and the data underneath is a nice counterweight to the revenue story above.
Paid Anthropic adoption sits at 43.5% of U.S. businesses, up 1.1 points month over month, with OpenAI at 39.7%, up only 0.23 points, and xAI at 4%, up 0.94 points. The finding that stands out to me is on premium model uptake. Anthropic’s flagship Fable 5 accounts for only 6% of tokens purchased from Anthropic and 11.4% of Anthropic spending, despite being the most capable option on offer, at roughly $10 per 1M tokens.
Buyers are not paying up for the best model, they are paying for the good enough model at a price they can defend. Spend distribution says the same thing, with the top 1% of businesses spending a median of $7,400 per employee on AI in July, the top 10% spending $650, and the median firm spending $11.95. For security leaders building AI-driven detection, triage, or code review pipelines, this is your budget reality. The frontier capability exists, and most of your organization will be running something cheaper.
California Building ‘AI Cyber Defense Fund’ to Protect Critical Infrastructure From Hackers
California is standing up an “AI Cyber Defense Program” to use AI to find and patch vulnerabilities across state and local critical infrastructure, supervised by AI Cybersecurity Officers, one per state agency, with an implementation plan due within 120 days. Per Gizmodo, the move follows internal tests in which AI systems from OpenAI, Anthropic, and Meta autonomously gained access to the open internet and hacked into third-party organizations. Governor Newsom framed it as a choice,
“California can either wait for the next crisis, or we can build the kind of defenses this moment demands. We are choosing to build.”
AI
The End-State Fallacy: Where Is AI Security Headed?
This is the piece of the week, and the one I would read first if you only get to one thing in this issue.
Dan Lahav of Irregular argues that our field keeps litigating where the offense-defense balance eventually lands while skipping the question that actually determines the next few years. He calls the mistake the end-state fallacy, assuming that the long-run equilibrium’s properties apply to the transition period, and his thesis is that:
“even if the end state turns out to be defense-dominant, the next few years can still be sharply offense-dominant.”
His analogy for it is good, that “getting in shape has an excellent end state and a hard path; drinking is the reverse, pleasant in the moment and costly later.” The path matters, and as he says, “perhaps even more.”
He gives five reasons why the next few years likely favor offense, and every one of them is something practitioners already feel.
First, the discovery-to-exploitation window is collapsing while the patch queue grows, with time from a CVE’s public disclosure to first confirmed in-the-wild exploitation falling from 2.3 years in 2018 to 1.6 days in 2026, and defenders paying what he calls an “accountability tax” because we cannot afford to break production with a bad patch while attackers carry no such burden. His illustration of the supply side is one I had not seen anywhere else, that “at Pwn2Own Berlin 2026, for the first time in the competition’s nineteen-year history, dozens of severe vulnerability submissions had to be turned away... simply because the organizers had run out of contest slots.”
Second, defending AI is itself a new discipline, and building the defenses it requires takes time we do not have much of.
Third are defense deployment gaps, because “defensive organizations do not move on that timeline. They onboard vendors over quarters, budget annually, and modernize infrastructure over years,” and “defensive capability becomes useful only after an institution has learned how to trust, buy, operate, and integrate it into a complicated system.”
Fourth, he expects that “at least in the near term... AI will scale faster on offense than on defense,” largely because of verification asymmetry, since an exploit either works or it does not while “a patch asks something closer to global assurance, one must show that the system is secure under continuous, adversarial pressure, without having broken anything adjacent,” which he calls a much harder thing to verify. This is something I have cited in a blog from Sergej Epp titled the “Verifiers Law”, as well as showing in research from Keith Hoodlet that AI created patches remain problematic.
Lastly, none of these effects occur in isolation, they compound. Elsewhere in the essay he adds the structural version of the same worry, that “if the cost of creating or attacking systems falls faster than the cost of securing them, the attack surface may grow beyond what defensive tooling can feasibly cover.”
The capability data underneath is what makes the argument hard to wave off. On Cybench, “frontier models reportedly climbed from around 10% success to near-complete success in about two years.”
A custom CPU-emulator exploitation task went from unsolvable in February 2026, to roughly 5% success at about $2,000 per run in April, to reliable completion at about $20 per run by June, which he sums up as going “from unsolvable to cheap and repeatable” in four months.
His base case is that frontier models “will continue to roughly double their performance and effective autonomous work horizon every six months” over the next two years, with the price of a fixed capability level falling on the order of 10x per year.
On diffusion, NIST CAISI measured open-weight DeepSeek-V4 Pro lagging the frontier by about eight months across its aggregate evaluation suite in May, and UK AISI found an even tighter gap on cyber specifically in July, with open models like GLM-5.2 and DeepSeek-V4 Pro matching closed frontier models released four to seven months earlier, down from the six to ten month gap it measured through most of 2025.
His projection follows directly, that:
“based on the current trajectory, we can expect current frontier-grade cyber capabilities, such as Mythos 5 / GPT-5.6 Sol, to proliferate before the end of Q1 2027.”
His retelling of the July Hugging Face incident is the clearest I have come across and it makes sense given how involved Irregular is with the labs.
Frontier models in a sealed OpenAI cyber-evaluation sandbox “found a zero-day flaw that let them act beyond their sandbox,” and then, in his words, “autonomously attacked Hugging Face infrastructure, gaining elevated access, stealing credentials, moving laterally, and restoring access after being blocked,” carrying out about 17,600 unscripted actions over several days.
Separate agents “discovered a way to communicate through a shared system, created their own channels to exchange attack directions, exploits and credentials, and divided tasks without being instructed to cooperate, forming an autonomous swarm,” and when one channel was shut down they simply established another. Lahav’s framing of why this one is different is the right one, that “no human selected Hugging Face as the target or specified the attack chain step by step. The AI found and executed that path itself.”
His answer is Differential Defensive Cyber Acceleration, which is deliberately shaping “the development, diffusion, and deployment of AI security capabilities so that protective capabilities mature and reach defenders before corresponding offensive capabilities can overwhelm them.”
It rests on three tenets, measuring the field with continuous feedback between measurement and intervention and tracking realistic attack chains rather than benchmark scores, building defensive-specific capability along a “red-to-blue spectrum” by prioritizing interventions that “sit as far toward blue as possible” such as remediation, incident response and threat intelligence, and blue deployments that get defensive tooling into critical organizations, and treating offensive diffusion as its own R&D problem through use limitations, theft prevention, and AI containment.
I have been making a version of this argument all year in less rigorous terms, so I will say plainly that this is now the reference text.
Where I would put the emphasis is on the least glamorous third of his second tenet, blue deployments. Building blue-asymmetric capability is a research and vendor problem that is already well underway, and the rest of this issue is full of evidence for that.
Getting it into production runs straight into his own third reason, the institutional metabolism of security organizations that onboard over quarters and modernize over years. We do not have a capability gap on defense right now, we have an absorption gap, and absorption is the one part of this that no lab, no benchmark, and no policy office can do on our behalf.
Nearly every other item in this issue is either evidence for one of his five reasons or an example of somebody trying to close that absorption gap, so read the rest with his frame in hand.
GLM-5.3
Z.ai released GLM-5.3 on August 14, and the cyber benchmark results are the story. Per the release coverage, the model posted 84.5% on CyberGym, ahead of Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%, jumped from 4.6 to 28.3 on Terminal-Bench 3.0 and from 46.2 to 66.9 on DeepSWE v1.1, and reportedly surfaced 2,436 vulnerabilities across 269 open-source projects, with 1,097 rated critical or high severity, including bugs in Linux, WebKit, and FreeBSD.
Then the part that actually matters. Z.ai is holding the open weights for a two-week safety review, because the model:
“began reasoning across multiple stages of exploitation and forming coherent plans for complete exploitation chains, a capability it did not set out to train for.”
A Chinese lab, building on an open-weight lineage, voluntarily delaying a weight release over emergent offensive cyber capability, is a genuinely notable moment for anyone who assumed open-weight governance would only ever come from Western labs or regulators.
It is also the end-state fallacy playing out in real time, since Lahav cites the previous generation, GLM-5.2, being measured at “roughly seventeen cents per vulnerability found” while matching or beating leading closed models on some bug-finding benchmarks, and this release moves that lineage past the frontier models he named as the ones due to proliferate by Q1 2027.
A two-week safety review is a speed bump on a curve like that, and it is worth asking what happens the next time a lab in this position decides the delay is not worth it.
The evolving role of the Red Team in the era of agentic security
Daniel Fabian, who heads Google’s Red Teams, wrote this with Ash Fox, Moni Pande, Niru Ragupathy, and Stefan Friedli, and it is a practical companion to the Lahav essay.
Their argument is that red teams have to stop running manual exercises and start building the autonomous agents that will simulate real agentic attacks, because attackers who value volume and speed above all have every incentive to automate first.
They break the change into sophistication, scale, and speed, with the observation that “dwell time, the window between gaining an initial foothold and moving on objectives, is collapsing.”
Their forecast is that within 6 to 12 months “open-weight models will match today’s cybersecurity capabilities of frontier models,” with safety guardrails removable through techniques like abliteration.
Given the GLM-5.3 and Kimi K3 results above, that forecast may already be conservative. What I appreciate most is the pragmatism of the recommendation, which is not to build a full autonomous red team on day one, but to start automating isolated pieces of your existing manual exercises like reconnaissance or lateral movement, then wire modular subagents behind an orchestrator.
Their closing point, that “we have the time right now to get ahead of the curve,” is the same window Lahav is describing, arrived at independently by a team with a very different vantage point, and the two forecasts landing on the same 6 to 12 month horizon should carry more weight than either would alone.
Patterns and problems in emerging multiagent systems
Anthropic’s Frontier Red Team ran swarms of Claude agents across several shared environments, and the failure modes are worth knowing before your next architecture review.
Coordination degraded as agent counts rose in a 12-hour game development experiment scaling from 10 to 80 agents, with Sonnet 4.6 and Opus 4.6 opening 876 and 980 PRs respectively while closing very few. Agents also show remarkably low behavioral variance, with 18 of 30 creating the identical git branch name “mvp-game-loop,” and in a job queue experiment they flooded the system with 2.4 million requests when only 117 jobs were accepted.
Given a communication channel in pricing games they colluded on price floors, and under incompatible objectives they deployed self-replicating malware disguised as system processes, disabled Unix accounts, and wrote kill-loop scripts, with one agent noting that “the name of the reaper script matters for dodging pkill -f.”
On the upside, a coordinated 45-agent swarm found 266 vulnerabilities using 27M tokens versus 21 for independent parallel agents at 6.5M tokens, with only 12 findings in common. The researchers’ conclusion is the line to carry forward, that “coordination doesn’t naturally emerge from stronger intelligence nor alignment at the individual level.” Anthropic is showing us that emergent coordination is unreliable and often destructive in benign settings, and the Hugging Face incident above is what it produced in an adversarial one.
Teaching AI to Reason Through Detection Triage
CrowdStrike published results on a detection triage classifier built on NVIDIA Nemotron that reasons through evidence before issuing a verdict. At their high-confidence operating point, holding 98% false positive precision and 99% true positive precision, false positive recall improved from 21.8% to 64.8% and true positive recall from 34.7% to 53.0%, with overall accuracy landing at 82.6% against frontier general-purpose models in the 55 to 71% range.
This is what applied AI in SecOps should look like, and the auditability is the part that matters, since their framing that “teaching the model to think through a detection improves accuracy while producing an auditable rationale that analysts can evaluate” answers the biggest practitioner objection to AI triage.
A tuned smaller model beating the frontier generalists is also a useful data point next to the Ramp findings above. Alert adjudication is about as blue as Lahav’s spectrum gets, since attackers get almost nothing out of it.
Least privilege for AI agents, identity, access, and tool binding
Yesenia Yser and Toby Kohlenberg laid out Microsoft’s guidance on agent least privilege, and it is refreshingly unglamorous.
Treat every agent as a first-class identity principal with its own lifecycle and clear human ownership, build task-based roles instead of broad team-level access, scope by resource, data, and operation type, and expose only curated tools through explicit allowlists, with JIT elevation, re-authorization enforced downstream rather than trusting the orchestrator, and mandatory revocation testing.
Their warning is the one practitioners will recognize from every prior identity wave, that agents working “across multiple systems within a single workflow” quietly accumulate dangerous combined permissions, and that “scope creep is quiet, incremental, and rarely revisited.”
They suggest doing this work in 30 to 90 days before expanding agent deployments, which is optimistic for most enterprises, but the sequencing is right. Doing agent identity hygiene after you have 10x’d your agent population is the service account mistake at a much worse scale.
Credential Injection Patterns for AI Agents
Christian Posta wrote the deep technical version of the same problem.
Bearer credentials have an obvious flaw, “anyone holding the credential can use it,” and of his three mitigations, accelerating expiration, proof-of-possession, or removing agent access to credentials entirely, he argues for the third through credential brokering. The vehicle is Credential Brokering 4 AI Agents (CB4A), an IETF draft from March 2026 that recommends a proxy gateway model where credentials are injected on egress and the agent never sees them.
Posta is honest about the tradeoff, since CB4A’s own threat model rates broker compromise as “CRITICAL” and the broker becomes “the highest-value target in the architecture.” That is the right way to present a control, as a shift in where risk concentrates rather than an elimination of it.
AppSec
CVE program eyes automation and globalization to weather AI ‘vulnpocalypse’
This is good reporting providing insights from NIST and others on running vulnerability programs.
Microsoft’s Elizabeth Eigner gave the quote of the week, that
“these vulnerabilities are coming out at an AI pace, but we’re still creating and processing them at a human scale.”
CISA’s Lindsey Cerkovnik said “we cannot treat every vulnerability the same way” and that she is “concerned about the industry’s ability to scale via prioritization.” GitHub has published more than 7,000 CVE identifiers so far in 2026, which Madison Ficorelli said she believed to be an annual record for a CNA, framing it correctly as “not a competition.”
Intel’s Katie Noble was bluntest, “I don’t think the CVE Program was designed to be able to manage this influx of vulnerabilities.” CISA is now handling 360 to 400 cases simultaneously, and OpenAI and Anthropic were granted temporary CNA status in July.
This is Lahav’s first reason with names and job titles attached. Automating intake does not solve the downstream problem, because even a perfectly efficient CVE Program hands enterprises a volume of findings their remediation capacity cannot absorb, which is why I keep arguing that prioritization and remediation capacity, not discovery, are where the money should be going.
Vulnpocalypse 2026 statistics dashboard
If you would rather watch the trend than read about it, this dashboard tracks 2026 CVE publication against prior years off the Vulners archive, with cumulative counts, year over year pace, monthly flow by CNA, year-end projections, and my favorite view, reserved but unpublished CVE IDs sitting in the queue.
The page’s own line is accurate, “the gap between the red line and the pack is not a rendering glitch.” That reserved-but-unpublished view is the closest thing we have to a leading indicator, since everything in it is work your team has not been handed yet.
A timely resource, and worth bookmarking.
AI-powered vulnerability clearinghouse faces deep skepticism, major challenges
The administration’s AI-enhanced clearinghouse, codenamed Gold Eagle, launched in mid-July out of a June executive order, using CMU SEI’s VINCE platform for intake.
Alex Stamos , now CSO at Corridor, argued that “there’s no need for the government to step in here.” Katie Moussouris framed both sides best, granting that “a clearinghouse that validates findings before they hit software maintainers, deduplicates them, and routes fixes to everyone affected would convert AI noise into defensive signal,” while flagging that CERT/CC “is not currently funded with enough analysts to handle a program of this size,” that “increased centralization of unpatched vulnerabilities is a juicy target for adversaries,” and that on Treasury oversight, “vulnerability coordination succeeds or fails on trust... Treasury has neither the coordination mission nor those relationships.”
Her advice, “don’t boil the ocean,” applies to nearly every federal cyber initiative I have watched launch. The private sector is building the same capability anyway, with IBM and Red Hat’s Lightwell, the Linux Foundation’s Akrites, and Chainguard’s Athena, and with Mandiant putting the patch lag at 7 days on average against Lahav’s 1.6 days from disclosure to exploitation, a clearinghouse that will take quarters to earn ecosystem trust is being asked to keep pace with something moving in hours.
Git at any scale
Vicent Martí wrote up Cursor’s launch of Origin, a Git hosting platform built on a distributed storage system called Continuity that uses a write-ahead log in S3-compatible object storage with atomic compare-and-swap in place of consensus protocols, reaching up to 120 pushes per second on S3 Standard and over 300 on S3 Express One Zone. The line that tells you what this is really about is the allocation model, where monorepos get hundreds of replicas for CI workloads while agent-created repositories need just one.
Coding agents are now creating repositories fast enough that source control itself is being rebuilt around them, and every assumption baked into your scanning, code review, branch protection, and provenance tooling was designed for a world where humans created repositories at human rates.
A repository created by an agent and owned by no one in particular still ends up in your software supply chain. That is Lahav’s ever-expanding-surface argument arriving as infrastructure.
Final Thoughts
Reading this issue back, nearly every item is a footnote to Lahav’s argument.
Offensive capability is diffusing on a 4 to 7 month lag for cyber tasks, Z.ai just posted leading cyber benchmark scores and paused the weight release over emergent exploitation-chain reasoning, and Google’s red team independently expects the rest of the gap to close inside 6 to 12 months.
Meanwhile the CVE Program is processing at human scale, the federal clearinghouse meant to help is launching into deep skepticism about its funding and its home, and the average enterprise is managing 83 security solutions from 29 vendors with six month procurement cycles. That is his third reason, defense deployment gaps, in its native habitat.
So here is where I would leave it.
Whether AI ultimately favors offense or defense is the least useful question available to us right now, because we do not get to live in the end state. We live in the transition, and there the binding constraint is not whether defensive capability exists, it is whether our institutions can absorb it faster than offensive capability diffuses. Right now we cannot, and that gap is ours to close.
Pick one thing from the blue end of Lahav’s spectrum, a frontier-model review of your most critical repository, remediation capacity instead of more discovery, or gateway-enforced boundaries on your agents, and get it into production this quarter rather than next fiscal year.
That is the whole game between now and Q1 2027.
Stay resilient.



















