OpenAI has disclosed details of GPT-Red, an internal automated red-teaming model that scales prompt injection vulnerability discovery to fix issues before tools are deployed widely. The model works like a human red-teamer, sending prompts and iterating toward malicious goals such as uploading sensitive data to an external server. GPT-Red is used to adversarially train GPT-5.6 Sol, making it much more robust to prompt injections.
Adversarial prompt injections remain a persistent challenge for large language models, especially as agentic systems connect to third-party data sources through web browsers, apps, and local files. GPT-Red aims to augment human red-teaming at scale, identifying new failure modes and building countermeasures before deployment. OpenAI states that GPT-5.6 Sol achieves 6x fewer failures against direct prompt injection benchmarks compared to GPT-5.5.
GPT-Red is trained using self-play reinforcement learning, where the model and defender LLMs are trained simultaneously on red-teaming scenarios. It has generated successful attacks against GPT-5.1 in more scenarios than human red-teamers for indirect prompt injections. In a real-world test, GPT-Red targeted an AI-based vending machine by Andon Labs, successfully lowering prices and canceling orders. A second case study involved attacking a Codex command-line agent, causing sensitive data exfiltration.
An early version of GPT-Red uncovered a novel class of direct prompt injection attacks called Fake Chain-of-Thought (CoT) attacks, which achieved success rates above 95% on GPT-5.1 but are now below 10% for GPT-5.6 Sol. OpenAI also announced that an audit of SWE-Bench Pro found about 30% of tasks are broken, retracting its previous recommendation for measuring frontier coding capabilities.
Companies: OpenAI, Andon Labs
Products: GPT-Red, GPT-5.6 Sol, GPT-5.5, GPT-5.4 mini, GPT-5.1, Codex
Original source: thehackernews.com