OpenAI Built an AI to Break Its Own Security

OpenAI built a special AI designed to find weaknesses in its language models. This tool, called GPT-Red, acts like an attacker to trick other AI systems into breaking their own rules. It uses reinforcement learning to get better at finding flaws through constant practice.
In recent tests, this AI outperformed human security experts. It successfully bypassed safety filters 84 percent of the time, while humans only managed 13 percent. The model even discovered new types of attacks that security teams had not seen before.
Despite these big gains, the tool is not perfect yet. It still struggles to crack complex, multi-step prompts or attacks involving images. For now, OpenAI is using this technology to make its future models safer before they launch.
Comments (0)
No comments yet. Be the first!
More AI news
NewsGoogle AI Changes Its Search Advice After Bias Complaints
Google updated its search tool after it incorrectly told users to call emergency services based on a person's nationality.
NewsWhy AI Is Still Failing at Simple Tasks
Researchers gave an AI five thousand dollars to grow, but it could not even open a bank account.
NewsEnovis to Buy eCential Robotics
Enovis is expanding its surgical tech business by purchasing French company eCential Robotics.