we_are_coded.by CODE · The world, decoded
БГ
OpenAI

OpenAI launches an automated attacker that hardens its models

OpenAIFrontier

OpenAI unveiled GPT-Red - an internal tool that automatically attacks its own models (red-teaming) and feeds what it learns back into their training. Per the company, the new version achieves noticeably fewer breaches under prompt injection than its predecessor from months ago.

In short
  • OpenAI announced GPT-Red - an automated red-teamer trained with self-play, built into the training process for future models.
  • Per OpenAI, it achieves around 6 times fewer successful breaches under direct prompt injection compared with a version from four months earlier.
  • The goal: security hardens automatically and continuously, instead of being done by hand for every new model.
Checked on16 July 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

How do you harden a model against abuse? You set people loose to attack it - so-called red-teaming - and plug the holes they find. It's slow and it doesn't scale. OpenAI announced it has automated the attacker itself.

The facts: on 15.07.2026 OpenAI unveiled GPT-Red - an automated red-teamer trained with self-play that attacks its models and feeds what it learns back into their training process. Per OpenAI, the new version achieves around 6 times fewer successful breaches under direct prompt injection compared with its predecessor from four months earlier. The numbers are the company's own and haven't been independently verified. Source: OpenAI, 'Unlocking Self-Improvement for Robustness'.

Pretty on paper - and probably genuinely useful, but the number is OpenAI's own and I take it with the usual reservation. More interesting than the '6 times' is the idea underneath: the defense learns by itself, in a loop, instead of people patching it for every new model. If it works, security stops being a snapshot taken before launch and becomes a process that keeps running. If it doesn't work, it's just another automation handing out false peace of mind.

An automated attacker turned on your own model is the right direction - but precisely because it's automated, nobody outside sees what it misses. You close the loop within yourself and ask the world to take your word for it. The number is good; independent verification is what will give it weight.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→OpenAI - Unlocking Self-Improvement for Robustness (GPT-Red)
Original: https://wearecoded.com/en/articles/openai-gpt-red-self-improvement.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news