OpenAI vs Anthropic: The Truth About ARC-AGI Scores

OpenAI recently announced that its new model, GPT-5.6 Sol, outperformed Anthropic’s Opus 5 on a difficult AI reasoning test. They reported a score of 38.3 percent. However, this high score was only achieved by using their own custom testing software.
When researchers used the official testing environment instead, the score for the OpenAI model dropped significantly to just 7.8 percent. By comparison, Opus 5 reached 30.2 percent using the standard rules without any special help.
This gap shows how much a test environment can change the results. It is a good reminder to look closely at the fine print before deciding which AI model is truly the best.
Comments (0)
No comments yet. Be the first!
More AI news
NewsNew AI Tool Fixes Mistakes in Financial Research
Researchers developed a new framework to stop AI from learning from its own bad data in financial trading models.
NewsPerplexity adds local AI to Mac apps
The new update lets your Mac handle some AI tasks directly on your computer.
NewsMeta Releases Muse Image Model on Fal
Meta has launched a new AI model on the Fal platform that plans and edits images on its own.