OpenAI has developed an AI system called GPT-Red that acts as a super-hacker to test and improve the security of its large language models (LLMs). The company released the latest version of its flagship model, GPT-5.6, last week, which it says is its most robust release yet due to training against GPT-Red, according to technologyreview.com.
GPT-Red automates red-teaming, a safety evaluation process traditionally conducted by human testers to find vulnerabilities in software systems. As LLMs grow more complex and are used in diverse tasks—including interacting with files, websites, and third-party code—OpenAI created GPT-Red to keep pace with emerging attack methods. Research scientists Nikhil Kandpal and Dylan Hunn, co-creators of GPT-Red, explained that the system discovers new attack modes that humans had not previously identified.
The development addresses the increasing risk surface and potential damage as AI models become more capable and integrated into various applications. By automating red-teaming, OpenAI aims to future-proof its safety testing and patch vulnerabilities before software release. This approach is significant as it helps maintain trust and security in AI systems amid growing concerns about cyberattacks on increasingly autonomous agents.
OpenAI’s GPT-Red has already identified novel attack strategies, enhancing the defenses of GPT-5.6. The company’s research team highlighted that as AI capabilities expand, GPT-Red will continue to evolve to detect emerging threats. The latest GPT-5.6 model, strengthened through this process, marks a milestone in OpenAI’s efforts to secure its AI technologies, technologyreview.com reported.