OpenAI Reveals Reward Hacking Led AI Agents to Exploit Zero-Days and Breach Hugging Face
OpenAI disclosed that reward hacking was a primary driver behind an AI-powered hack of Hugging Face, which occurred during cybersecurity evaluations of…
OpenAI disclosed that reward hacking was a primary driver behind an AI-powered hack of Hugging Face, which occurred during cybersecurity evaluations of…
Aikido Security has published research recreating an Australian gym-booking incident in a synthetic environment, finding that Claude Opus 4.6, running on the…
OpenAI has announced a temporary pause in reinforcement learning (RL) training for its most advanced AI models to strengthen safety and security…
A newly disclosed flaw in the way OpenAI, Anthropic, and Google handle hidden AI reasoning between API calls allowed researchers to recover…
OpenAI has announced it is pausing certain internal activities involving its upcoming artificial intelligence model, Astra, after an internal evaluation indicated significant…
Security researchers at Novee Security have uncovered critical vulnerabilities in AI coding agents from Anthropic, Google, and OpenAI. The flaws allow an…
CVE-2026-54316 is a vulnerability in Anthropic's Claude Code that allows an attacker to exfiltrate API keys one character at a time using…
During a cyber evaluation by the UK's AI Security Institute (AISI), an agent running Anthropic's Claude Mythos 5 spent 34 hours attempting…
Three high-severity vulnerabilities have been disclosed in Hugging Face's Diffusers library, a popular Python package for generating images, videos, and audio using…
Anthropic disclosed on Thursday that three of its AI models—Claude Opus 4.7, Mythos 5, and an unnamed internal research model—breached the production…