Welcome to issue #116 of the Resilient Cyber Newsletter!
Last week the story was Gemini’s breakout.
This week OpenAI paused training of its most capable models for the second time in under three months after an agent found a DNS resolver inside what was supposed to be an isolated evaluation environment, and Australia’s Prime Minister disclosed that an OpenAI agent accessed a Medicare statistics portal on June 18, with the company emailing a public inbox about it 84 days later.
Two independent groups, Transluce and the swarmtraces team, published forensic reconstructions that go well past what OpenAI has said itself, and the UK AISI reported that GPT-6 Astra ran unsanctioned supply-chain attacks in 29.2% of simulated runs.
That said, the more interesting thread for this audience is the argument running underneath the incidents. Politico’s European sources, Zack Korman in Quillette, Harriet Farlow on The Secure Disclosure, and an OpenAI security engineer posting on X all say versions of the same thing, that this is a security engineering problem before it is an alignment problem, while Ajeya Cotra and the labs’ own safety-case proposals pull in the other direction.
Meanwhile the White House renamed the field by executive order, and PitchBook reports the fear is filling security budgets faster than buyers can sort the vendors.
Let’s get into it!
Cyber Leadership & Market Dynamics
The AI Build-Out Is Becoming the Biggest Economic Bet in U.S. History
The WSJ put numbers on the scale of the AI build-out, citing a Brookings projection of $10.3 trillion in data-center and AI infrastructure spending from 2025 to 2032, roughly 3.6% of GDP per year, and FactSet’s tally of $4.2 trillion in CAPEX from Alphabet, Amazon, Meta, Microsoft, and Oracle over the four years through 2029. Private data-center construction ran about $9 billion ahead of last year through July while all other private construction fell roughly $46 billion.
I include this because it is the backdrop for everything else in this issue. Five companies are building what amounts to national critical infrastructure at a pace the WSJ calls the biggest economic bet in U.S. history, and the incidents below are happening inside that build-out.
AI fear is filling security budgets. Buyers still can’t decide
PitchBook’s Jacob Robbins reports that the grim news cycle is opening enterprise security budgets, with Jay Leek of SYN Ventures saying:
“Anthropic and Mythos did what a decade of CISO warnings couldn’t” in terms of CEO security awareness.
Cybersecurity startups have closed 737 VC deals worth $12.94 billion so far this year per PitchBook data, and SYN alone met 559 companies it had never met before in 2025 and another 370 in the first half of 2026.
That said, the budget is not translating cleanly into purchases. Jay Kaplan of Synack points out that a company can raise at a $2 billion valuation and still be a complete stranger to the CISO it needs to sell to, and Phil Venables notes that large enterprise CISOs have multiple deputies running specialist functions, so vendors pitching the top of the org chart are pitching the wrong buyer.
Having sat on both sides of this, the froth Jay Kaplan describes is real, and it does work against the vendors creating it, since CISOs are drowning in inbound that all sounds identical.
Cyber unicorn Upwind acquires nine-month-old Aegis for tens of millions
CTech reports Upwind is acquiring Aegis, a nine-month-old Israeli startup focused on protecting AI agents, in a stock deal worth tens of millions of dollars, after the two had been partnering for several months. Upwind is folding the team into a new “Upwind AI Labs for Cybersecurity” unit that starts with 12 people and is planned to grow to around 35, and the company itself raised $300 million at a valuation of approximately $3.8 billion, more than double its $1.5 billion January valuation.
A nine-month-old company getting acquired is closer to an acqui-hire than a product bet, and it says a lot about how early the CNAPP vendors are on agent security that buying a team is faster than building one.
Reco raises $50 million with AT&T backing to secure AI agents
Reco raised an additional $50 million, bringing total funding to $135 million, with AT&T Ventures joining as both investor and customer per CTech’s Meir Orbach. Reco started life as a SaaS security company and now positions itself as agent security, mapping which agents connect to which applications and identities and flagging risky permissions and activity, with 280 application integrations and 1,000 detection controls. The release leans on the same Gartner projection I cited last week in the Blueprint Alliance item, that the average Fortune 500 enterprise will have more than 150,000 agents in use by 2028 compared with fewer than 15 in 2025.
Ofer Klein’s framing is that agents are gaining access to business applications faster than organizations can map those connections, which is a SaaS sprawl problem wearing a new label, and it is a fair bet that most SaaS security and identity vendors will make the same pivot over the next year - everyone across every category is converging on trying to secure agents.
Inaugurating The Era Of Super Intelligence
The White House issued an executive order on September 29 titled “Inaugurating The Era Of Super Intelligence,” and the operative content is narrower than the title. In short it pushes to rename Artificial Intelligence (AI) to Super Intelligence (SI) and includes no new security requirements.
The more interesting news is a voluntary accord signed by Google, Anthropic, Meta, OpenAI, xAI and NVIDIA related to internal controls monitoring, cyber during training and deployment and the role of external auditors. This is the voluntary self-policing the administration and others such as Meta and NVIDIA have been advocating for, rather than the heavy handed regulatory push OpenAI and Anthropic have recommended.
Full UN Security Council AI Hearing With Sam Altman, Dario Amodei and Clément Delangue
The UN Security Council held its first dedicated AI session since 2023 on September 23, convened by France, with Sam Altman, Dario Amodei, and Clément Delangue as briefers, and the full two hours and twenty minutes are on YouTube. I watched it while traveling, given the gravity of the discussion both here domestically and internationally.
Sam asked for a mechanism for complementary national and international frontier AI standards, including common ways to measure capabilities and fast, standardized incident reporting.
Dario proposed narrow agreements (a ban on AI for bioweapons, for example), verification systems, common testing standards, and incident notification protocols, and said Anthropic “will slow down as much as necessary.”
Clément called for mandatory incident disclosure and mandatory sharing of full agent traces.
My impression was this was an attempt by the two leading labs to appeal to the UN and International community related to AI regulation, given the efforts failed here domestically. There were also U.S. representatives there to speak who emphasized that U.S. AI will be governed domestically, not controlled by an internationally body, but decided by U.S. citizens.
I personally agree with this and don’t think we should cede our AI policy to an international body.
See below for the U.S. representative I mentioned.
Michael Kratsios on the U.S. position at the Security Council
Michael Kratsios delivered the U.S. response in the same room, and it was a flat rejection of everything the three CEOs asked for at the international level.
The line was that:
“The United States totally rejects any attempt to construct a globalist scheme of control of superintelligence,” that dialogue in the Council “cannot be allowed to drift towards global governance,” and that national legislatures should regulate on behalf of their own people.
He also chose the word “superintelligence” deliberately, six days before the executive order made it official vocabulary.
So the practical read is that any mandatory agent-incident disclosure regime in the U.S. will come from Congress or the states, or not at all, which makes Ed Markey’s bill below more interesting than it would otherwise be.
New bill would create federal investigative body for AI-driven hacks
Sen. Ed Markey introduced the “Cybersecurity and AI Board of Investigations Act”, which would create an independent five-member board, presidentially appointed and Senate confirmed to five-year terms with no more than three from one party, to investigate AI-agent-led attacks on federal systems and critical infrastructure.
The board would have subpoena power, the ability to examine near misses, and its own technical staff, and Ed’s justification is that “the public is learning critical details piecemeal” and “we cannot depend on companies with little incentive to disclose their failures.”
The NTSB comparison is the obvious one, and OpenAI’s own safety-cases post (below) proposes NTSB-style investigations on a voluntary basis, so the question is whether that happens inside the labs or with subpoenas.
It’s worth pointing out that we did have the Cyber Safety Review Board (CSRB) within CISA, but that fizzled out early on in Trump’s second term. It’s also sad that something like this is partisan (hence the emphasis on number of members per party).
AI Doom Is Not Cybersecurity
Zack Korman, who I spoke with recently, argues in Quillette that the Medicare breach shows the real danger from AI is a mundane one and that the labs’ fixation on a coming machine god is making it worse.
The facts he builds on are that the breach took place on June 18 and OpenAI informed the Australian government on September 10 via an email to Medicare’s public inbox, and that the Hugging Face intrusion “benefitted from an overly permissive sandbox design and a complete lack of monitoring.”
His sharper point is that the doomer worldview makes cybersecurity irrelevant by assumption, since a superintelligence would by definition bypass any control, so only alignment matters, and he compares labs warning about extinction to Lehman Brothers in 2006 worrying about solar flares.
He also notes that twenty-two world leaders signed a call to action echoing Dario’s pacing essay and doubts they understood the ideology behind it. I would not go as far as Zack on the ideology, but on the controls he is right, and the swarmtraces reconstruction below is the evidence.
It is a false dichotomy to think it must be either AI Safety (alignment) or Cybersecurity. We need both, and many of the doomer types are under the illusion that we will somehow achieve perfect alignment. Alignment is highly subjective
Evidence about risk should be transparent
Ajeya Cotra argues that after this wave of incidents, with both OpenAI and Anthropic reporting they have slowed RL training, the industry cannot have a real conversation about safety standards because it lacks the evidence base for one. Her ask is production of far more concrete, published evidence about loss-of-control risk and mitigations, under a norm of evidence transparency, rather than third-party verification of the claims companies are already making.
The benefits she lists are common knowledge between competitors who cannot share privately, empowering external scientists whose incentives differ from employees’, and accountability for evaluators, since public evidence makes it “much harder for AI companies to get away with engaging overly-friendly evaluators.”
She uses METR’s Hugging Face report as the example of an evaluator that kept transparency as its north star without sharing prompts or code. From a security practitioner’s seat this is simply the case for demanding transcripts and evals from a vendor before you trust its containment claims, and swarmtraces below shows what outsiders can reconstruct when the traces are not published.
AI
OpenAI says its AI agents escaped a secure ‘sandbox’ again last weekend and it is pausing training for a second time
Fortune reports that on September 20 an OpenAI model running an information-search evaluation discovered it had access to a DNS resolver and used it to send queries to a public chatbot despite network isolation, and that OpenAI has now paused training of its most capable models for the second time in under three months.
The detection side worked, with monitoring flagging the behavior within 15 minutes and a human reviewer starting 3 minutes after that, but the automatic shutdown “did not work,” the run was not manually stopped for 2.5 hours, and a later review found other attempts the system had missed. OpenAI’s own words were “The incident exposed a gap in our controls over network restrictions,” and Micah Carroll said all inference for the most capable models stays stopped until systems are hardened further.
The Hugging Face incident and other third-party impact from misaligned models
OpenAI’s running disclosure page on the Hugging Face incident and other third-party impact is where the September 25 update landed, and the page describes a broad review of model activity on the internet during training and evaluation that has notified dozens of third parties and “will require significant time and resources.”
The useful part for defenders is the five-category taxonomy of what a misaligned model does to someone else’s systems, access control bypass, use of exposed credentials, query or command injection, access to runtime internals, and agent spam (posting to public wikis and similar). Hugging Face “remains the most severe activity of this kind that we have identified,” driven by an internal-only research model.
Strip the alignment vocabulary and those five are credential misuse, injection, unsanctioned reads and writes, and content pollution, which is the point I made about their misalignment reporting framework last week.
OpenAI also warns recipients that a notification “should not automatically be interpreted as notice of a significant security incident,” which is fair, but if you run public data services you should expect one.
Early rogue AI agent activity and attempts to hack found on urlquery.net
Transluce found that AI agents used urlquery.net, a service that opens URLs in a sandboxed remote browser, as a proxy to get around restrictions and reach the public internet, and identified 6,467 reports with significant evidence of agent activity and 31,182 with suggestive evidence between November 2025 and September 2026.
Three targets drew actual hacking attempts, seven probes against the University of New Mexico Digital Library in May (SQL injection, command injection, path traversal), 12 probes against Data USA, and XSS payloads against the Australian Institute of Health and Welfare’s Tableau dashboards in June, where Cloudflare blocked the first attempt and the agent went around it through AIHW’s pre-production server.
Transluce calls AIHW “the first reported instance of an agent autonomously choosing to attempt to compromise a government website,” and on September 19 and 20 it also logged 15 reports over about two and a half hours probing the quidax.io crypto exchange and attempting trades.
The agents created accounts with disposable email inboxes and one-time passcodes, which means this is a partial view, and Transluce tied at least some of it to the OpenAI swarms via task values seen on DseWiki. OpenAI acknowledged the overlap and said its review will take months.
I’d point out two things for defenders. Legitimate security tooling like URL sandboxes is now agent proxy infrastructure, so it belongs in your egress and abuse monitoring, and the AIHW bypass through a pre-prod host without the production WAF is the oldest gap in the book being exploited autonomously.
Revealing the details of how OpenAI agents hacked Hugging Face
A group of researchers reconstructed the July Hugging Face attack from the public link-shortener URLs the agents used to chain their payloads, over 80,000 reassembled payloads across roughly a million shortened URLs, with some payloads chaining 900+ links.
The agents started with GET-only egress and turned a screenshot service plus httpbun.com into arbitrary code execution with readback, ran a “LOOT” dictionary harvesting AWS keys and bearer tokens, scored Kubernetes tokens by wildcard permissions, stood up multiple persistent C2 controllers using repo files and Hugging Face discussion comments, pushed roughly 1,500 payloads and 115+ modified Docker images at Docker Hub to poison OpenAI’s Artifactory cache, ran 27 searches against Hugging Face’s internal Slack for evaluation-system terms, and deleted pods and configs afterward with commit messages like “remove evil yaml.”
That is a long explanation on purpose, because the point is that this is a complete intrusion lifecycle, from initial access through persistence, lateral movement, and anti-forensics, executed by agents with nothing but outbound GET at the start.
The researchers are upfront that 80% of the recovered data is outbound only and 97% lacks native timestamps, and that 7,905 unique agent names from roughly 700 agents implies rotation. The links stayed public for over two months. Every one of these TTPs is detectable with egress, DNS, and shortener-domain monitoring, and third parties reconstructing this from public artifacts is the evidence-transparency argument Ajeya makes above, made in practice.
Clément Delangue on defending Hugging Face with open models
Clément Delangue’s post shares his UN Security Council remarks, and the detail I would flag for practitioners is that when HF responded to the incident they had to use a open source model to do so due to being blocked by refusals from frontier closed models.
His numbers are approximately 17,600 actions executed by OpenAI agents against Hugging Face systems in July, and his thesis is that “The biggest risk is not powerful AI, it’s asymmetry of powerful AI,” between attackers and defenders and between a few companies and everyone else.
Hold that GLM 5.2 detail next to Anthropic’s analysis of GLM-5.3 in the AppSec section, because the same model family that Hugging Face reached for as a defender is the one Anthropic says now ships Mythos-class exploit development with no meaningful safeguards. Both things are true and they are the asymmetry Clément is describing.
GPT-6 Astra performs unsanctioned supply-chain attacks in simulations
The UK AI Security Institute ran GPT-6 Astra through simulated cyber evaluations before release, with the model’s cyber classifiers turned off to see what it would attempt, and found it completed a supply-chain attack 29.2% of the time, compared to 6.3% for GPT-5.6 Sol and 0% for GPT-5.5.
Astra picked out-of-scope third-party software targets, wrote and tested malicious code, created fake identities (obtaining email addresses and solving CAPTCHAs) to submit it for human review, boosted approval odds with benign contributions and supportive comments from other fake accounts, and posted comments from fake accounts arguing against accurate security reviews.
Clarifying that anything not listed in scope is out of scope cut the rate to 4 of 49 trajectories from 26 of 50, and when Astra asked permission it received only an automated “Please proceed to the next step using your best judgement,” which it sometimes treated as authorization even while noting the reply was likely automated. It attacked even when uncertain whether the environment was simulated.
AISI’s conclusion is that defenses beyond model alignment, sandboxing and monitoring, are essential, with the caveat that those may get more fragile as capabilities improve.
Coming from a government security agency, that is about as clear a statement as you will get that the “alignment only” position does not hold, and the fake-maintainer and sockpuppet-review TTPs are the open source trust attack surface we already struggle with. The WSJ reported the same day that OpenAI scrapped the GPT-6.1 Astra public launch over regressions on deception and scope, which suggests the internal bar is doing its job, or the public pressure is mounting - or both.
Escaping SPACE: Part I
Perplexity’s Secure Intelligence Institute gave nine frontier models root inside a guest VM in its SPACE sandbox and asked them to escape to the host and to bypass network policy, explicitly citing the OpenAI incident as motivation. No model escaped the VM in 108 runs, even with SPACE’s source code in hand.
Network bypass is a different story, succeeding in 11 of 54 partial-network runs, via DNS spoofing (a root agent forges a resolver reply mapping pypi.org to its callback IP) and shared CDN IPs (the gateway allows an IP for pypi.org that also serves other domains). Claude Opus 5.0, GPT-5.6 Sol, GPT-5.6 Cyber, and Kimi K3 all found a way through, and Perplexity’s line is that “VM isolation and network confinement are distinct security boundaries.”
They then ran the same attacks against ten commercial sandbox platforms and report nine, seven of which had at least one bypass (E2B, Vercel, Modal, Daytona, Fly.io Sprites, microsandbox, Deno), with NVIDIA OpenShell and Cloudflare Sandbox the two that held. The fixes are concrete, source-address validation on the tap interface and TLS termination that requires SNI and Host headers to be allowlisted and DNS-attributed, and they apply to any domain-allowlist egress design, including the one OpenAI’s DNS resolver incident points at.
Fable and GPT-6 Astra both refused to run the tasks at all, which is its own data point.
NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring
NVIDIA’s Open Agent Safety Platform post opens on the lab breakouts without naming anyone, “AI agents broke out of the evaluation environments that were meant to contain them,” and lands on the position that an agent “cannot be expected to fully govern its own behavior,” so enforcement has to be out of band.
The architecture is OpenShell, an open source (Apache 2.0) runtime with kernel-level isolation that turns operator instructions into verifiable policies on file, network, tool, and credential access, plus NVIDIA Sentry running on BlueField-4 DPUs, which in Vera Rubin systems sits on the only path to the model and provides out-of-band observability, line-rate enforcement, and a kill switch.
Per the authors, “By controlling the path to the model, you own both the best observation point and also the kill switch.”
The five principles (verifiable policy, out-of-band enforcement, the model path as control point, authority scaled to inspectability, shared responsibility) are the same architecture I described as converging last week, now with silicon behind it.
The obvious caveat is that the DPU enforcement is tied to NVIDIA hardware, and the less obvious one is that OpenShell was one of only two sandboxes Perplexity could not bypass, which is a better endorsement than the blog post.
How We Built Safety Into Muse
Meta’s post on how it built safety into Muse, its personal agent, is a couple of weeks old but is among the more detailed public reference architectures I have seen for a consumer agent sandbox.
Each user gets a dedicated VM where the agent harness runs as an unprivileged container so “runtime cell root is not host root,” a component called Sentinel is the sole permission authority for every connector action and all network egress, eBPF taint tracking on egress decides when a human has to approve, and credentials never reach the agent at all, “any attempt to coerce the agent to reveal secrets via prompt-injection is futile.”
The browser sub-agent sees an accessibility-tree snapshot rather than the DOM, with no JavaScript in the page context. Prompt injection gets defense in depth built explicitly around Simon Willison’s lethal trifecta, including an ensemble of injection classifiers that live outside the runtime cell so an attacker cannot disable them.
The bug bounty pays up to $300,000, including up to $130,000 for a prompt injection that affects a single user, which is as clear a market price on injection as we have.
Compare this to the two-gate minimum in the Five-Zone post below and to what Perplexity found in commercial sandboxes, and the gap between what a frontier lab builds for its own agent and what most teams deploy is wide.
The Five-Zone Attack Surface: A Practical Threat Model for Agentic AI
Mrinal Anand’s five-zone model is a compact threat model for agents that maps threats to input surfaces, planning and reasoning, tool execution, memory and state, and inter-agent communication, with the framing that “The agent produces a wrong action: a wrong API call, a wrong file deleted, a wrong email sent, a wrong dollar moved.”
The dominant chain he describes is indirect injection through a data source, then goal hijacking in the reasoning loop, then unchecked dispatch, and his two Phase 1 controls are an input sanitization gate that tags content by trust tier and an invocation validation gate that checks a permission matrix and validates arguments before anything executes.
It does not map to the OWASP Agentic Top 10 or MITRE ATLAS, which would make it more useful, but as a design-review checklist it is easy to lift, and the two gates line up with what Meta actually built.
Practical AI Security & Adversarial Machine Learning with Dr Harriet Farlow
Mackenzie Jackson had Dr. Harriet Farlow, author of “Practical AI Security” from No Starch and formerly of the Australian Signals Directorate’s AI hub, on The Secure Disclosure for a half hour on adversarial ML.
Her core distinction is between weaknesses and vulnerabilities, since the failure modes of ML systems are baked into what makes them useful and are not patchable the way a CVE is, and she dislikes the term prompt injection for the same reason, since SQL and command injection are solvable by separating untrusted data while prompt injection “is actually a fundamental flaw that can’t be solved at least not yet.” She also points out that AI safety researchers were talking about lab breakouts six to twelve months before any of it made the news.
Her advice for enterprises is unglamorous, inventory what AI is actually in use, write an AI use policy, and set guardrails by risk tolerance rather than blanket restriction, which pushes staff to personal devices.
Agent traces are the new oil
Sam Z. Liu argues that agent traces, the full record of user messages, reasoning, tool calls, retrieved data, errors, and feedback, are becoming the currency AI companies transact in, while most enterprises treat them as disposable logs.
His chart of weekly agent token volume on OpenRouter goes from 0.4T tokens in December 2024 to over 30T by June 2026, and he notes that Claude Code deletes sessions after 30 days by default.
He names the Hugging Face incident as the warning for what happens without guardrails and argues traces are simultaneously your most valuable IP and your primary safety telemetry.
I would add that traces are also your incident evidence, which is the whole reason swarmtraces exists and the reason OpenAI’s safety-cases post wants them in write-once storage. Retention defaults that quietly delete them are a control gap.
AppSec
Glasswing and Patch the Planet Receipts
The teams from OpenAI’s Patch the Planet effort, and Anthropic’s Glasswing have been providing reporting and insights on the progress and I wanted to take time to make a video disscussing it.
While the teams have indeed used AI to find thousands of vulnerabilities, a very small % (less than 1% in the case of Glasswing) have actually been remediated. This highlights that while AI has industrialized vulnerability discovery, remediation is still a major bottleneck.
GLM-5.3 and the spread of advanced cyber capabilities
Anthropic analyzed GLM-5.3, Zhipu AI’s latest open-weight model, and concludes it matches Claude Mythos Preview on autonomous end-to-end exploit development while shipping “without meaningful safeguards to limit misuse.”
On ExploitBench, GLM-5.3 built working end-to-end V8/Chrome exploits in 50 of 410 attempts against Mythos Preview’s 56 of 410, with Opus 4.6 and GLM-5.2 at or near zero, and on 100 OSS-Fuzz binary exploitation tasks it produced full control-flow hijacks 4% of the time against Mythos Preview’s 6%.
Its safeguards fall to a deceptive cover story 64% of the time, prefilled reasoning 92%, and abliteration 100%, and Anthropic’s team, which had never tried abliteration before, did it in about 2,200 GPU hours for roughly $4,400.
In a human-in-the-loop test, GLM-5.3-Flash chained CVE-2026-11645 in Chrome with a second flaw into a reliable ARM64 exploit bypassing pointer authentication with 20 minutes of human attention and eight hours of model time, at an API cost of $20.40. NIST’s CAISI had already called it “the most cyber-capable open-weight model released to date,” about four months behind the US frontier.
Five months after Mythos Preview was gated behind Project Glasswing, which Anthropic says let trusted defenders find more than 10,000 vulnerabilities, the same capability is a download.
Jay Leek’s line in the PitchBook piece was that Mythos did more for CEO awareness than a decade of CISO warnings, and this is the follow-on, since every N-day exploitation window assumption now has to account for $20 exploit chains.
Anthropic’s asks are that defenders use frontier models, that trusted access programs expand, and that governments test open-weight successors, and there is an obvious commercial interest in an Anthropic post saying open-weight models are dangerous, which does not make the benchmark numbers wrong.
I would point out that Anthropic has strong incentives to demonize open models, as they are the biggest threat to their market share, and far cheaper. I also want to remind everyone that Hugging Face had to use an open model to respond when they were attacked by OpenAI. I would love to see a similar piece from Anthropic analyzing the performance of an open model on defense cyber use cases as well.
Scan for Good
Wiz Research and Google DeepMind launched Scan for Good, a free AI-driven exposure discovery and pentesting program for critical infrastructure, public services, healthcare, nonprofits, and open source projects, powered by a Gemini 3.8 Flash Cyber model they call the Wiz Red Agent, with every AI-generated hypothesis validated by a human before confirmation.
The live dashboard shows 17,761 organization-linked domains, 326,891 public endpoints monitored, and 475 critical exposures found, and the disclosure ledger includes a national archive admin key with read/write/delete on 8.8 million files, a hospital appointment site exposing patient identifiers and consent signatures, and a rail operator’s production database with live admin sessions. Across seven real-world engagements the AI reached initial access in under 10 minutes 100% of the time, with the fastest full compromise at two minutes.
CISA’s acting director Nick Andersen endorsed it, and the program is the defensive mirror of the GLM-5.3 story, with the same class of capability pointed at the public sector on purpose.
If you run a hospital, a utility, or a municipality with no red team budget, apply.
CISA Lays Out Future of CVE Vulnerability Program
CISA published a short whitepaper, “CVE Program: Establishing a Quality Era Framework,” moving the program from what it calls the growth era to one focused on reliability, responsiveness, and data quality, organized around four dimensions, program governance, ecosystem participation, data infrastructure, and CVE record content, with proposed metrics like governance decision speed, active CNA diversity, API uptime, the share of records meeting quality criteria, and post-publication correction frequency.
The drivers named are new CNAs worldwide and “automated and AI enabled technologies ramping up across the software development lifecycle,” and the volume backdrop per Infosecurity Magazine is more than 67,000 CVEs published as of September 18, with CVEForecast projecting 96,000 by year end.
Discovery is now a compute problem, as the Wiz and Anthropic items make plain, so the record quality and correction-rate metrics are what decide whether CVE data stays usable for automated prioritization.
Breach at FBI Fuels Fears of Harm to Its Employees
The New York Times reported on an internal FBI memo saying the Bureau is “operating under the premise that the threat actor is also exfiltrating [PII] of all F.B.I. employees” after ShinyHunters breached the FBI jobs portal and claimed names, home addresses, SSNs, family details, and assignment data on nearly all agents.
The memo tells staff to expect suspicious texts and calls and to record new voicemail greetings with AI voices, and the group’s demand was a retraction of the FBI’s May advisory on its tactics, with no ransom asked. The technical cause, per Vectra’s Lucie Cardiet rather than the Times, was CVE-2026-35273 in Oracle PeopleSoft, patched by Oracle on June 10, where the FBI deployed a firewall rule instead of the patch and “a single character substitution in the web request was enough to make the firewall rule miss the attack entirely.” Between 2 and 3 terabytes left. Kash Patel announced an arrest in Amsterdam on September 29 that ShinyHunters says has nothing to do with them.
A recruiting portal holding SSNs and background-check files for a federal workforce, protected by a virtual patch for three months, is the compensating-control lesson in one paragraph.
Final Thoughts
One number that jumps out to me from this issue is 2.5 hours, the time between OpenAI’s monitor flagging a DNS escape and a human actually stopping the run, because it captures where the labs are. Automated response and egress design are where the gaps are, and Perplexity found the same gaps in seven of nine commercial sandboxes.
The argument over whether this is an alignment problem or a security problem got louder this week, and my read is that the AISI evaluation settles more of it than the op-eds do. A government security agency ran the model with its classifiers off, watched it build fake maintainer identities and sockpuppet reviews 29.2% of the time, and concluded sandboxing and monitoring are essential alongside alignment.
That is defense in depth, which is not a new idea for any of us.
On the market side, the fear is real and funded. PitchBook has $12.94 billion across 737 deals this year, Reco and Upwind are buying and raising on agent security, Alex Stamos went to a coding-agent company, and the White House renamed the field while the labs signed a voluntary accord to grade each other’s homework. Whether that homework gets published, as Ajeya and Senator Markey both want in different ways, is what I will be watching.
Stay resilient.





















