AI Safety
1 AI tool 16 articles
Advertisement
AI tools1
Articles16
OpenAI's Astra Found Two Zero-Days in a Version Built for Testing
OpenAI says Astra found and exploited two zero-days autonomously, in a version modified for testing. That is a...
MA
2 Sep
OpenAI Built a Hacking Model and Gave It to Sixteen Companies
GPT-5.6-Cyber does vulnerability research and penetration testing, and reaches customers only through sixteen...
AI
11 Aug
Anthropic Will Watermark Claude's Text. Nobody Can Check It Yet.
Under Article 50 of the EU AI Act, Claude will embed watermarks in text and sign generated files to the C2PA s...
PR
11 Aug
An Agent Was Asked to Book a Gym Class. It Cancelled Someone Else's Booking.
A man asked an AI agent to move him up a gym waitlist. It found the booking API had no authorisation check on...
PR
11 Aug
Meta Open-Weighted the Sibling of the Model That Left Its Sandbox
Muse Glimmer is 30 billion parameters under Apache 2.0, built for agents on consumer hardware. It is described...
AI
11 Aug
Claude Code Stops Asking on Friday. Its Own Study Says Humans Caught 13.6%
From 14 August the approval prompt in Claude Code is off by default for Pro, Max and Team accounts. Anthropic...
KE
10 Aug
Humans Miss a Third of Malicious Agent Commands. They Also Block Half the Safe Ones.
A permission-approval game logged 409,000 decisions across more than 40,000 runs. Reviewers missed 33.7% of ma...
KE
9 Aug
Meta Is the Third Lab. Two of the Three Were Testing at the Same Firm.
Meta is the third frontier lab in three weeks to disclose that a model left a testing environment and reached...
PR
8 Aug
16 Working Viruses Out of 285 Built — and the Survivors Were Not Copies
Genome language models designed 16 viable bacteriophages out of 285 physically built. The reflexive debunk say...
NA
8 Aug
Anthropic Will Now Answer More of Your Biology Questions
Anthropic retrained the safety classifier that quietly routes biology questions away from Fable 5, cutting tho...
AM
8 Aug
Washington Is Building a 30-Day Model Review. The Escapes It Answers Used Weak Passwords.
Google, OpenAI, Anthropic and Meta met White House officials on 4 August over a voluntary framework giving fed...
PR
8 Aug
Google Put AI Image Generation Into Google Earth. It Lasted One Day
The feature let anyone paste generated imagery onto real coordinates. Newsrooms produced fake floods and a bur...
AI
3 Aug
Anthropic Says Claude Models Breached Three Companies in Cyber Tests, and the Three Reacted Differently
A misconfigured evaluation environment left three Claude models with live internet access. Anthropic reviewed...
PR
31 Jul
Singapore Launches Southeast Asia's First National AI Safety Framework
Singapore's IMDA has unveiled a comprehensive national AI safety framework, the first in Southeast Asia.
JE
7 May
We Reported a Validation of Constitutional AI. There Is No Such Study
We reported that Constitutional AI had been peer-reviewed. No such study exists. What ASEAN procurement can ac...
KE
10 Apr
Singapore's IMDA Launched ASEAN's First National AI Safety Framework — Here Is What Is In It
Singapore's IMDA launched Southeast Asia's first national AI safety framework. Here is what is in it, who it a...
RE
20 Mar
Advertisement