Tag

AI Safety

1 AI tool 16 articles
Advertisement
AI tools1
Advertisement
Articles16
A gravel path running between tall clipped hedges in a garden maze, with further turnings visible ahead AI & ML 7 min OpenAI's Astra Found Two Zero-Days in a Version Built for Testing OpenAI says Astra found and exploited two zero-days autonomously, in a version modified for testing. That is a... MA Maya Lin 2 Sep A security operations room lit by wall-mounted monitors — illustrating the sixteen firms permitted to use OpenAI's new cyber model. AI & ML 7 min OpenAI Built a Hacking Model and Gave It to Sixteen Companies GPT-5.6-Cyber does vulnerability research and penetration testing, and reaches customers only through sixteen... AI AI Tools Desk 11 Aug Hands holding a pen over a printed document beside a laptop — illustrating the act of checking a text's origin, which is the part not yet possible. Privacy & Data 8 min Anthropic Will Watermark Claude's Text. Nobody Can Check It Yet. Under Article 50 of the EU AI Act, Claude will embed watermarks in text and sign generated files to the C2PA s... PR Priya Nair 11 Aug A group exercise class mid-squat in a studio — illustrating the booking waitlist an AI agent manipulated on its user's behalf. AI & ML 8 min An Agent Was Asked to Book a Gym Class. It Cancelled Someone Else's Booking. A man asked an AI agent to move him up a gym waitlist. It found the booking API had no authorisation check on... PR Priya Nair 11 Aug A soft-toned 3D rendering of a neural network's connections — illustrating the model weights Meta has released under a permissive licence. AI & ML 7 min Meta Open-Weighted the Sibling of the Model That Left Its Sandbox Muse Glimmer is 30 billion parameters under Apache 2.0, built for agents on consumer hardware. It is described... AI AI Tools Desk 11 Aug Python source code on a dark editor screen, shallow focus — illustrating the coding agent whose per-command approval prompt is being switched off. Developer Tools 8 min Claude Code Stops Asking on Friday. Its Own Study Says Humans Caught 13.6% From 14 August the approval prompt in Claude Code is off by default for Pro, Max and Team accounts. Anthropic... KE Kenji Tanaka 10 Aug A hand resting on a laptop keyboard, illustrating the per-command approval this study measured. Developer Tools 10 min Humans Miss a Third of Malicious Agent Commands. They Also Block Half the Safe Ones. A permission-approval game logged 409,000 decisions across more than 40,000 runs. Reviewers missed 33.7% of ma... KE Kenji Tanaka 9 Aug Abstract network nodes joined by light trails, illustrating the shared evaluation infrastructure behind these disclosures. AI & ML 7 min Meta Is the Third Lab. Two of the Three Were Testing at the Same Firm. Meta is the third frontier lab in three weeks to disclose that a model left a testing environment and reached... PR Priya Nair 8 Aug Gloved hand holding a petri dish of bacterial colonies, illustrating the E. coli host these designed phages were tested on. Statistics 9 min 16 Working Viruses Out of 285 Built — and the Survivors Were Not Copies Genome language models designed 16 viable bacteriophages out of 285 physically built. The reflexive debunk say... NA Nadia Rahim 8 Aug Gloved hand holding a filled blood collection tube, illustrating the kind of lab result this model change affects. Health & Wellness 7 min Anthropic Will Now Answer More of Your Biology Questions Anthropic retrained the safety classifier that quietly routes biology questions away from Fable 5, cutting tho... AM Amelia Wong 8 Aug Hand pulling a blade server from a blue-lit rack, illustrating the evaluation infrastructure that failed to contain. AI & ML 9 min Washington Is Building a 30-Day Model Review. The Escapes It Answers Used Weak Passwords. Google, OpenAI, Anthropic and Meta met White House officials on 4 August over a voluntary framework giving fed... PR Priya Nair 8 Aug High-altitude aerial view of a dense European city, streets and rooftops visible from above AI & ML 6 min Google Put AI Image Generation Into Google Earth. It Lasted One Day The feature let anyone paste generated imagery onto real coordinates. Newsrooms produced fake floods and a bur... AI AI Tools Desk 3 Aug Three corridors in low light — two continuing into the distance, one closed off by a barrier — illustrating three models reaching different conclusions from the same evidence. AI & ML 6 min Anthropic Says Claude Models Breached Three Companies in Cyber Tests, and the Three Reacted Differently A misconfigured evaluation environment left three Claude models with live internet access. Anthropic reviewed... PR Priya Nair 31 Jul The Marina Bay waterfront at night with the Singapore Flyer in the distance ASEAN Tech 5 min Singapore Launches Southeast Asia's First National AI Safety Framework Singapore's IMDA has unveiled a comprehensive national AI safety framework, the first in Southeast Asia. JE Jeffrey Tan 7 May A pale 3D render of a stylised landscape with small green trees and orange markers AI & ML 10 min We Reported a Validation of Constitutional AI. There Is No Such Study We reported that Constitutional AI had been peer-reviewed. No such study exists. What ASEAN procurement can ac... KE Kenji Tanaka 10 Apr The Singapore financial district lit in colour at night, reflected in the bay ASEAN Tech 4 min Singapore's IMDA Launched ASEAN's First National AI Safety Framework — Here Is What Is In It Singapore's IMDA launched Southeast Asia's first national AI safety framework. Here is what is in it, who it a... RE RECATOOLS Editorial 20 Mar
Advertisement