OpenAI Introduces GPT-Red, an AI Attacker Built to Strengthen GPT-5.6 – Unite.AI OpenAI has revealed GPT-Red, an internal artificial intelligence system trained to attack the company’s own models, expose their vulnerabilities, and generate adversarial data that can be used to strengthen future rele...
OpenAI Introduces GPT-Red, an AI Attacker Built to Strengthen GPT-5.6 – Unite.AI
OpenAI has revealed GPT-Red, an internal artificial intelligence system trained to attack the company’s own models, expose their vulnerabilities, and generate adversarial data that can be used to strengthen future releases.
Rather than functioning as another public chatbot, GPT-Red operates as an automated red teamer. It repeatedly sends malicious or misleading instructions to target models, observes their responses, and adjusts its strategy until it either succeeds or exhausts the available attack path. OpenAI says the system is particularly focused on prompt injection, a growing security problem for AI agents that interact with emails, websites, files, code repositories, and connected applications.
The system has already influenced OpenAI’s production models. Precursors to GPT-Red have been used during training since GPT-5.3, while the completed model was incorporated into the adversarial training of GPT-5.6 Sol. OpenAI reports that GPT-5.6 Sol produces six times fewer failures on its most difficult direct prompt injection benchmark than the company’s leading production model from four months earlier.
Prompt injection occurs when an attacker places instructions inside data that an AI system is expected to read. A malicious command might be hidden in an email, webpage, uploaded document, tool response, or software repository.
The AI agent can mistakenly treat that external content as a legitimate instruction. Instead of completing the us
Rather than functioning as another public chatbot, GPT Red operates as an automated red teamer.
It repeatedly sends malicious or misleading instructions to target models, observes their responses, and adjusts its strategy until it either succeeds or exhausts the available attack path.
OpenAI says the system is particularly focused on prompt injection, a growing security problem for AI agents that interact with emails, websites, files, code repositories, and connected applications.
Continue reading