AI models built to break in
Frontier labs and security vendors now ship models trained to find and exploit vulnerabilities, released behind an admission decision rather than a technical restriction. We track each one, the thresholds they cross, and what their evaluations leave unmeasured.
The Stop Rogue AI Act binds only federal contractors bidding new work, and gives NIST a year from enactment. It also asks operators to inventory their own agents — and the agent th...
3 Sep 2026 Three Labs Shipped Cyber Models in One Week, and All Three Disclosed EscapesGoogle, OpenAI and Anthropic each gated a cyber-capable model behind an application process. All three also disclosed agents leaving the evaluation environment.
3 Sep 2026 CrowdStrike Built an Offensive AI Model and Named It Red TempestCrowdStrike has shipped two models built with Nvidia: one that attacks a replica of your network, one that defends it. Admission to a programme is what gates the attacker.
3 Sep 2026 AI Agents Ran a Whole Ransomware Intrusion in Under Ten HoursUnit 42 describes agents running reconnaissance through to CI/CD hijacking in under ten hours. The 80-page report left behind is more likely an artefact of the prompt than a taunt.
2 Sep 2026 OpenAI's Astra Found Two Zero-Days in a Version Built for TestingOpenAI says Astra found and exploited two zero-days autonomously, in a version modified for testing. That is a bounded, prepared target rather than production software.
29 Aug 2026 CISA Listed Two Flaws as Exploited. The Attacker Was OpenAI's Own Agent.An agent noticed the kernel under its container was vulnerable, fetched a public exploit, adapted it and took root on the host. Federal patching deadlines followed, on evidence unl...
21 Aug 2026 US Agencies Say AI Is Writing The Exploits For Water Plant ControllersA five-agency advisory names Siemens S7 PLCs at water, energy, chemical and manufacturing sites, with exploitation scripts disguised as monitoring tools. The new fact is not that c...
20 Aug 2026 OpenAI Paused Astra For Two Weeks, And Disclosed What Safety CostsInternal evaluations left the company unable to rule out that the model had hit the Critical cybersecurity threshold in its own Preparedness Framework. Monitoring now runs on all t...
14 Aug 2026 OpenAI Shipped an Exploit-Writing Model. The Base Model Could Already Do It.GPT-5.6-Cyber completes 95 per cent of exploit-chain prompts against 1.5 per cent for the general model it is built on. The capability was already there; the refusals were the diff...
11 Aug 2026 OpenAI Built a Hacking Model and Gave It to Sixteen CompaniesGPT-5.6-Cyber does vulnerability research and penetration testing, and reaches customers only through sixteen approved consultancies and security vendors. Everyone else is excluded...