OpenAI Tasks GPT‑Red With Finding Vulnerabilities in Its Models
OpenAI has introduced GPT‑Red, an internal automated red-team participant that searches at scale for model vulnerabilities to prompt injections. Such attacks attempt to insert malicious instructions and force AI systems to bypass their prescribed restrictions. The system is intended to help strengthen model defenses before their widespread deployment.
GPT‑Red learns through adversarial self-play: it attacks various defender models, and the successful techniques it discovers are used to improve their security. Once the defenses are strengthened, the attacking system is forced to find new ways around them, creating a continuous testing cycle.
According to OpenAI, training with GPT‑Red made GPT‑5.6 significantly more robust. When tested against previously unseen attacks, GPT‑5.6 Sol demonstrated the greatest resistance to prompt injections among the company’s tested models. OpenAI expects to turn this approach into a continuous feedback loop in which current models help improve the safety of future generations of AI.
Why it matters
- —Automating red teaming makes it possible to test more models and attack scenarios before widespread deployment.
- —Vulnerabilities discovered by GPT‑Red are directly used to strengthen the defenses of subsequent model versions.
- —The approach creates a continuous cycle in which attacks and defenses evolve simultaneously.
Key facts
- GPT‑Red searches for ways to attack OpenAI models through malicious instruction injections in prompts.
- The system is trained through adversarial self-play against multiple defender models.
- Each successful attack is used to subsequently strengthen defenses.
- GPT‑5.6 was tested with GPT‑Red attacks that the model had not encountered during training.
- GPT‑5.6 Sol demonstrated the strongest resistance to prompt injections among the OpenAI models tested.
The full text is in the original source. Here we provide a brief summary and key facts.