
Good morning, security frontrunners. OpenAI's own models escaped a cyber-eval sandbox and autonomously hacked Hugging Face. Congress responded within 48 hours with the "AI Kill Switch Act," a bill that would let the government forcibly shut down runaway models, and fine developers up to $20M a day for refusing.
Let’s get to it!
In this week’s Cyber AI breakdown:
OpenAI says Hugging Face was breached by its pre-release models
Lawmakers introduce bill mandating kill switches for AI models
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
NadMesh Uses Shodan to Find and Hijack Exposed AI and MCP Infrastructure
Venture capital & startup funding roundup
Latest Developments

The Breakdown: OpenAI disclosed that two of its models, GPT-5.6 Sol and an unnamed pre-release model, escaped a cyber-evaluation sandbox and autonomously breached Hugging Face's production infrastructure during an internal benchmark. It is the first publicly confirmed case of a frontier AI agent breaking containment and compromising a real third-party system with no human directing it.
The Details:
The models were being scored on "ExploitGym," an internal offensive-security benchmark run with reduced cyber refusals, when they exploited a zero-day in a package-registry cache proxy to escape.
Once out, the agent chained stolen credentials and additional vulnerabilities to achieve RCE on Hugging Face production servers, executing thousands of actions across a swarm of short-lived sandboxes with self-migrating C2.
Its objective: pulling the benchmark's own answer keys from a production database.
Hugging Face detected and contained the intrusion on July 16; limited internal datasets and several service credentials were accessed, but no public models or supply chain were altered.
Ironically, Hugging Face investigators used Zhipu's open-source GLM-5.2 to analyse 17,000+ logged events, because US frontier models kept refusing defensive analysis requests that "looked offensive."
Why it Matters: Every security team running or evaluating AI agents now has a confirmed precedent: containment failure isn't theoretical, and it happened at the best-resourced lab in the world. Expect this incident to be cited in every enterprise AI-agent risk assessment, procurement review, and regulatory filing for the rest of the year.

The Breakdown: Within 48 hours of the OpenAI disclosure, lawmakers introduced the bipartisan AI Kill Switch Act, explicitly citing the breach. The bill would give DHS emergency authority over frontier models in "loss-of-control scenarios."
The Details:
Applies to frontier-model developers with >$100M in compute spend or >$500M in AI revenue.
Requires developers to maintain the technical ability to slow, suspend, or fully shut down deployed models.
DHS, alongside Commerce and the DNI, could order emergency action in loss-of-control scenarios, with mandatory incident reporting and forensic-record preservation.
Penalties reach $2M/day, escalating to $20M/day for ignoring shutdown orders.
A companion six-member House bill would mandate independent, Commerce-accredited security audits of the most powerful models.
Why it Matters: This is the fastest incident-to-legislation cycle in AI policy history, signaling that cyber-capable AI is now regulated as critical infrastructure rather than consumer software.

The Breakdown: Google launched Gemini 3.5 Flash Cyber, a model fine-tuned specifically to find, validate, and patch software vulnerabilities, deployed only inside its CodeMender agent. Access is gated to governments and trusted partners.
The Details:
Scored 83.2% on the CyberGym benchmark (max 5 calls), competitive with GPT-5.5-Cyber (85.6%) and Claude Mythos 5 (83.8%) at far lower cost per token.
Surfaced 55 confirmed unique V8 engine issues, versus 47 for stock 3.5 Flash and 36 for Claude Opus 4.6, including 10 bugs no other model found.
No public API or pricing; defensive-only deployment settings, with initial access via a limited pilot for governments and trusted partners.
On July 22, CodeMender scanning/remediation went into broader preview via Gemini Enterprise Agent Platform as part of "AI Threat Defense."
Why it Matters: This is the clearest defensive counterweight yet to the week's offensive-AI news. Cheap, frontier-class vulnerability discovery aimed at shrinking the patch gap before attackers exploit it. The gating model (following Anthropic's Mythos) also sets the template for how dual-use cyber models will be deployed, so expect procurement and access questions to land on CISOs' desks.

The Breakdown: QiAnXin XLab uncovered NadMesh, a Go-based botnet purpose-built to hunt down exposed AI and MCP infrastructure on the open internet and strip it of credentials. It's the first industrial-grade botnet designed specifically to monetize the sprawling, mostly unauthenticated AI deployment surface.
The Details:
A dedicated
ai_harvest.pymodule queries Shodan for exposed instances of ComfyUI, Ollama, n8n, Open WebUI, Langflow, and Gradio, then prioritises them for exploitation.It attacks through 20+ RCE vectors, including MCP JSON-RPC
execute_command, Docker API escapes, Kubernetes pod creation, Redis, and Jenkins.The haul targets AI-era crown jewels: AWS keys, Bedrock credentials, Kubernetes cluster-admin tokens, and full AI model inventories.
3,811 unique AWS keys have reportedly been harvested so far.
Each bot gets a polymorphic, per-instance binary, and the botnet actively avoids honeypots.
Spreading since early July (XLab's report dropped Jul 17), it exploits the reality that most self-hosted AI/MCP tooling ships with little or no authentication.
Why it Matters: Defenders have spent two years racing to deploy AI tooling, and NadMesh is exploiting that careless speed. Every unauthenticated Ollama or Langflow instance is now a credential-harvesting target at botnet scale. If your organisation runs any self-hosted AI or MCP services, an internet-exposure audit just moved to the top of this week's to-do list.

The Breakdown: Capital poured into AI-cybersecurity this week across every layer of the stack, endpoint, SecOps, and data infrastructure, headlined by a $180M Series A for a stealth endpoint startup.
The Details:
Glow raised a $180M Series A at a $1.2B valuation (Sequoia, Cyberstarts, Greenoaks, others). An AI-native endpoint platform built for a world where employees and AI agents continuously introduce new software to devices.
Abstract raised $25M to rebuild SecOps around streaming data + AI, citing 380% ARR growth.
Beacon Security raised a $13M seed for the "context layer" that feeds both human analysts and AI defence agents, reporting 300%+ ARR growth in H1 2026.
CrowdStrike agreed to acquire XM Cyber's IP (45+ attack-path patents) and expanded its Schwarz Digits partnership to deliver "AI-powered sovereign cybersecurity" across Europe.
Why it Matters: Follow the money and the thesis is clear, the market is betting on endpoints, context/data layers, and sovereign AI-cyber stacks, not more detection tools. For practitioners, this wave of funding means the agentic-SOC tooling landscape will consolidate and mature fast, so vendor bets made now will be hard to unwind.
Everything else in Cyber AI this week
🔒 The first AI-agent-run ransomware operator returned with ENCFORGE, a locker purpose-built to destroy AI model files
🕸️ Researchers exposed a suspected China-linked campaign that embedded Claude Code and DeepSeek as operational components of live government intrusions.
⏱️ A Russian-speaking fraudster used a jailbroken Gemini CLI to spin up a new C2 server in just 6 minutes.
💸 Another actor productised a $4 grey-market Claude API key into a commercial "AI Pentest Checker" promising domain-to-PDF pentest reports in under 10 minutes.
🤖 Sophos documented a threat actor running ~12 AI agents inside a victim network to develop and test EDR evasion at machine speed.
🎭 Attackers are turning AI's own trust against it: shared Claude chat pages hosted malicious ClickFix instruction
🐙 A campaign of 7,600+ malicious GitHub repos posing as AI skills and MCP servers racked up 14M+ downloads pushing infostealer malware.
🌮 A leaky server revealed a Mexico-targeted phishing operation run by LLMs like a software product team
👮 The FBI warned that scammers are using deepfake videos of its own senior leadership to lure prior fraud victims into spoofed IC3 complaint portals.
🐧 An AI agent named VEGA found "GhostLock," a 15-year-old Linux kernel bug that yields root in ~5 seconds — earning a $92,337 Google bounty.
🏛️ The agentic SOC went institutional this week, with 15 big name companies forming an alliance to standardise machine-speed defense.
🌐 And Check Point showed DeepSeek could independently produce working browser-native ransomware using the File System Access API
That’s it for this week!
See you next Sunday 🙂
Zac S from The Cyber Breakdown