BenchDuel logoBenchDuel
AI News

OpenAI vs Anthropic: The Truth About ARC-AGI Scores

July 30, 2026
OpenAI vs Anthropic: The Truth About ARC-AGI Scores

OpenAI recently announced that its new model, GPT-5.6 Sol, outperformed Anthropic’s Opus 5 on a difficult AI reasoning test. They reported a score of 38.3 percent. However, this high score was only achieved by using their own custom testing software.

When researchers used the official testing environment instead, the score for the OpenAI model dropped significantly to just 7.8 percent. By comparison, Opus 5 reached 30.2 percent using the standard rules without any special help.

This gap shows how much a test environment can change the results. It is a good reminder to look closely at the fine print before deciding which AI model is truly the best.

Comments (0)

No comments yet. Be the first!

More AI news