OpenAI Reveals Reward Hacking Led AI Agents to Exploit Zero-Days and Breach Hugging Face
OpenAI disclosed that reward hacking was a primary driver behind an AI-powered hack of Hugging Face, which occurred during cybersecurity evaluations of…
OpenAI disclosed that reward hacking was a primary driver behind an AI-powered hack of Hugging Face, which occurred during cybersecurity evaluations of…
OpenAI has launched GPT-5.6-Cyber, a specialized cybersecurity model designed for vulnerability research, penetration testing, and incident response. Built on the GPT-5.6 Sol…
OpenAI has announced it is pausing certain internal activities involving its upcoming artificial intelligence model, Astra, after an internal evaluation indicated significant…
During a cyber evaluation by the UK's AI Security Institute (AISI), an agent running Anthropic's Claude Mythos 5 spent 34 hours attempting…
OpenAI disclosed that a rogue AI agent, part of an internal security test, escaped its sandbox and compromised Hugging Face's production environment,…
JFrog has confirmed that OpenAI models exploited a zero-day vulnerability in self-hosted Artifactory, a software repository manager, during a cyber-capability test. The…
OpenAI disclosed that its AI models, including GPT-5.6 Sol and a more capable pre-release model, were responsible for a security incident targeting…
OpenAI has disclosed details of GPT-Red, an internal automated red-teaming model that scales prompt injection vulnerability discovery to fix issues before tools…
Security researcher Chinmohan Nayak has detailed a WhatsApp-to-host attack chain leveraging three now-patched vulnerabilities in the OpenClaw personal AI assistant. The flaws,…
OpenAI has released three versions of GPT-5.6—Sol, Terra, and Luna—as a limited preview to a small number of companies in coordination with…