OpenAI vs Anthropic: The Truth About ARC-AGI Scores

OpenAI recently announced that its new model, GPT-5.6 Sol, outperformed Anthropic’s Opus 5 on a difficult AI reasoning test. They reported a score of 38.3 percent. However, this high score was only achieved by using their own custom testing software.
When researchers used the official testing environment instead, the score for the OpenAI model dropped significantly to just 7.8 percent. By comparison, Opus 5 reached 30.2 percent using the standard rules without any special help.
This gap shows how much a test environment can change the results. It is a good reminder to look closely at the fine print before deciding which AI model is truly the best.
Comments (0)
No comments yet. Be the first!
More AI news
NewsTencent Releases AngelSpec to Speed Up AI Models
Tencent just released a new open-source tool that makes large AI models run much faster.
NewsCut Your Claude PDF Costs by 99% With Token Saver
A new open-source tool lets you chat with PDFs in Claude while saving almost all your token costs.
NewsMoonshot AI Releases New Tool for Faster AI Training
Moonshot AI just released MoonEP, a new library designed to speed up the training of large AI models.