
Happy Sunday, security frontrunners. AI-powered bug hunting is on pace to double 2025’s record, as Google turns Big Sleep into a fully autonomous V8 pipeline and Microsoft pushes the same shift into enterprise defence with Project Perception.
Let’s get to it!
In this week’s Cyber AI breakdown:
AI-Driven Vulnerability Discovery Is on Pace to Double 2025’s Record
Microsoft Introduces MAI-Cyber-1-Flash and Project Perception
7AI Launches Federated SIEM and a Builder for Custom Security Agents
Tenable Hexa AI Becomes an Always-On Exposure-Management Fleet
watchTowr Project Red Reproduces New Vulnerabilities Autonomously
Latest Developments

The Breakdown: AI-powered vulnerability research is pushing the number of disclosed software flaws towards twice the record set in 2025. Defensive discovery currently appears to be outpacing confirmed exploitation, but organisations are receiving substantially more vulnerabilities to assess and patch within shorter exploitation windows.
The Details:
The US National Vulnerability Database recorded 45,207 vulnerabilities by 27 July, already approaching the total recorded throughout 2025.
Oracle patched 1,449 vulnerabilities in July, compared with 309 in July 2025.
Microsoft disclosed 642 vulnerabilities, nearly five times its July 2025 count.
Chrome fixed 433 vulnerabilities, compared with 11 in the equivalent 2025 update.
CISA’s Known Exploited Vulnerabilities catalogue has not recorded a comparable increase, but average time-to-exploitation reportedly fell from 72 hours in 2025 to 24 hours in 2026.
Why it Matters: Vulnerability management is becoming a throughput problem as AI generates more valid findings than traditional triage, testing and deployment processes can comfortably absorb. Cyber teams will need stronger asset context, exploitability analysis and automated patch validation to distinguish urgent exposures from a rapidly expanding backlog.

The Breakdown: Microsoft’s Project Perception is an agentic security system that continuously identifies attack paths, evaluates their risk and coordinates corrective actions across an organisation’s environment. Its first specialised model, MAI-Cyber-1-Flash, is being deployed inside Microsoft’s MDASH multi-agent vulnerability-discovery system.
The Details:
Project Perception coordinates red agents for attack-path discovery, blue agents for investigation and green agents for remediation.
MAI-Cyber-1-Flash initially targets software vulnerability management as part of the existing MDASH agent team.
Microsoft reports that the new configuration achieved 96% on CyberGym, 12 percentage points above Mythos, although the result has not been independently reproduced.
Microsoft claims MAI-Cyber-1-Flash can reduce operating costs by almost 50% compared with the current MDASH configuration.
The system uses shared, near-real-time context across identities, endpoints, applications, data and cloud systems, with public preview scheduled for 3 August.
Why it Matters: Perception moves Microsoft’s security strategy beyond AI-generated recommendations towards a closed loop in which agents discover, investigate and remediate exposures together. Its integration across Microsoft’s security ecosystem could make agentic exposure management accessible to large numbers of enterprises, but will also make governance and approval boundaries increasingly important.

The Breakdown: 7AI’s Federated SIEM allows autonomous security agents to investigate and act on telemetry without requiring every data source to be copied into a central platform. The accompanying 7AI Build capability lets organisations encode their own workflows, institutional knowledge and investigative techniques into reusable agent skills.
The Details:
Federated SIEM queries data across existing SIEMs, data lakes, cloud platforms and 7AI’s own storage.
Detection is separated from storage, allowing organisations to retain data where it already resides and potentially reduce duplicate ingestion costs.
7AI Build supports custom agentic workflows, investigative skills and complete AI-native security services.
Each customer receives a context graph connecting users, assets, policies, applications, historical investigations and organisational knowledge.
7AI says its agents have completed more than nine million investigations and returned over one million analyst hours, although the figures are vendor-reported.
Why it Matters: Federated investigation could weaken the assumption that every security event must be centralised before an SOC can analyse it effectively. Custom agent skills could also preserve experienced analysts’ investigative knowledge, but organisations will need careful testing to ensure that local workflows do not automate incorrect assumptions.

The Breakdown: Tenable is extending Hexa AI into a persistent fleet of agents that continuously identifies, prioritises and helps remediate exposures. Unlike isolated assistant sessions, the agents retain environmental context and execute scheduled workflows across vulnerability, asset and identity data.
The Details:
Hexa AI agents maintain context about an organisation’s environment between individual exposure-management tasks.
Daily routines can include scanning, vulnerability triage, remediation orchestration and verification that patches removed the exposure.
Weekly routines cover configuration drift, identity exposures, ticket progress and executive or compliance reporting.
Workflows are grounded in Tenable’s asset, vulnerability, identity and attack-path data and can incorporate built-in, third-party or customer-created agents.
The fleet is expected to become available during August for Tenable One Foundation and Advanced customers; independent accuracy and remediation benchmarks have not been published.
Why it Matters: Persistent agents could turn exposure management from a periodic assessment activity into a continuously operating control loop. The critical question for security teams will be how much remediation authority to delegate before mistakes, incomplete context or unsafe changes create operational risk.

The Breakdown: Project Red uses AI to analyse emerging vulnerability information and automatically construct a safe, working reproduction of the flaw. These reproductions feed watchTowr’s exposure-validation and mitigation systems, reducing the time between public disclosure and defensive action.
The Details:
watchTowr says Project Red reproduced CVE-2026-63030, or wp2shell, in 22 minutes and before exploitation was observed in the wild.
The system analyses vendor advisories, source-code changes, product updates and patches to determine the vulnerable behaviour.
It builds a non-destructive check that proves exploitability rather than relying only on product or version matching.
Reproductions are tested against patched and unpatched environments, processed through hardened CI/CD and reviewed by a human researcher.
Project Red supplies working checks to watchTowr’s Rapid Reaction exposure-validation and Active Defense mitigation capabilities.
Why it Matters: Autonomous reproduction helps defenders answer the operationally important question of whether a vulnerability is genuinely exploitable before weaponisation becomes widespread. It could significantly shorten exposure-validation times, although safe testing and human review remain essential when automatically generated exploit logic reaches production systems.
Everything else in Cyber AI this week
🧪 UK government testing found Kimi K3 stronger than GLM-5.2 on exploit-development benchmarks but substantially behind leading US models on arbitrary code execution.
🏛️ An unattended Hermes AI agent was observed enumerating systems during an intrusion targeting Thailand’s Ministry of Finance.
📧 The Russian-supported Laundry Bear group used AI-assisted code in a zero-click Zimbra campaign exposed by cyber agencies from 15 countries.
🔓 A penetration test uncovered a critical authorization failure in a Claude-built financial application processing identity and payment information.
🗺️ Stream Security launched StreamForce, allowing security agents to inherit a continuously updated map of cloud, identity, SaaS and on-premises assets.
🛡️ Cobalt’s new autonomous pentest promises findings within 24 hours while retaining a human pentester to direct scope and execution.
⚔️ Assail argues that the agent harness—including memory, orchestration and governance—determines offensive capability more than the underlying language model.
🦋 SonicWall has joined Project Glasswing to use Claude Mythos 5 for vulnerability discovery and code review across its software portfolio.
🛌 Google has turned Big Sleep into a fully automated pipeline that discovers, reproduces and validates vulnerabilities in Chrome’s V8 engine.
That’s it for this week!
See you next Sunday 🙂
Zac S from The Cyber Breakdown