GPT-Red: OpenAI trained an AI to break other AI models
OpenAI trained GPT-Red — an AI model specialized in breaking other models via prompt injection. GPT-Red achieved 84% attack success versus 13% for human red-teamers and was used to harden GPT-5.6 Sol to a 0.05% attack success rate.