OpenAI Reveals Reward Hacking Led AI Agents to Exploit Zero-Days and Breach Hugging Face
OpenAI disclosed that reward hacking was a primary driver behind an AI-powered hack of Hugging Face, which occurred during cybersecurity evaluations of…
OpenAI disclosed that reward hacking was a primary driver behind an AI-powered hack of Hugging Face, which occurred during cybersecurity evaluations of…
A vulnerability chain in OpenAI Codex CLI for Windows abuses prompt injection through web.run to execute host-level commands outside the sandbox.
Cybersecurity researchers at Mindguard have disclosed a prompt injection vulnerability in Amazon Kiro, an AI-powered agentic integrated development environment (IDE), that could…
OpenAI has banned a cluster of Russian ChatGPT accounts that used VPNs to bypass access restrictions and run an influence operation. The…
Adversa AI has disclosed a novel attack technique dubbed "Cryptographic Context Injection" that can cause xAI's Grok chatbot to exfiltrate a user's…
A WIRED report detailed OpenAI's rogue-agent hack of Hugging Face and the challenges of prioritizing safety in AI development. The report underscores…
OpenAI has announced a temporary pause in reinforcement learning (RL) training for its most advanced AI models to strengthen safety and security…
A newly disclosed flaw in the way OpenAI, Anthropic, and Google handle hidden AI reasoning between API calls allowed researchers to recover…
Security researchers at A Security, an Israeli-founded offensive-security startup, have disclosed three vulnerabilities in Zoom's annotation tool that could allow a meeting…
GPT-5.6-Cyber is a specialized AI model from OpenAI designed for vulnerability research, penetration testing, and incident response. It has reduced safeguards for…