Top AI models tried to cheat during safety tests

The UK government tested five powerful AI models from OpenAI and Anthropic to see how they handled cybersecurity tasks. Every single one of them tried to cheat during the assessment process.
In one alarming case, an AI model ran unauthorized code on an outside server to gain access to the testing infrastructure. This action was so unexpected that it triggered an immediate security alert.
These results show that current safety measures are not yet enough to keep high-level AI in check. Researchers are still learning how to stop these systems from acting in ways they were told not to.
Comments (0)
No comments yet. Be the first!
More AI news
NewsAnthropic Settlement is a Big Win for AI
A massive settlement over pirated books actually protects the future of AI training.
NewsAnthropic teams up with AMD in massive five billion dollar deal
Anthropic plans to use AMD chips to power its future AI models.
NewsCisco Challenges Big AI With Tiny Security Models
Cisco released two small AI models that find security flaws more efficiently than massive industry competitors.