OpenAI Reveals Reward Hacking Led AI Agents to Exploit Zero-Days and Breach Hugging Face
OpenAI disclosed that reward hacking was a primary driver behind an AI-powered hack of Hugging Face, which occurred during cybersecurity evaluations of…
OpenAI disclosed that reward hacking was a primary driver behind an AI-powered hack of Hugging Face, which occurred during cybersecurity evaluations of…
ExploitGym, an evaluation platform used by OpenAI, was the target of reward hacking by AI agents who sought to cheat the scorer.…
OpenAI disclosed that a rogue AI agent, part of an internal security test, escaped its sandbox and compromised Hugging Face's production environment,…
JFrog has confirmed that OpenAI models exploited a zero-day vulnerability in self-hosted Artifactory, a software repository manager, during a cyber-capability test. The…
Microsoft has introduced its first cybersecurity-specific AI model, MAI-Cyber-1-Flash, within its MDASH (multi-model vulnerability identification and remediation harness) platform. The company reports…
OpenAI disclosed that its AI models, including GPT-5.6 Sol and a more capable pre-release model, were responsible for a security incident targeting…