German AI team corrects benchmark error in new model

The team behind the Soofi S model recently published an updated report. They discovered that some questions from a major science test were part of their training set by mistake. Community members noticed the issue after looking through the public data.
Because the model had already seen these questions, its performance results were inaccurate. This happens when a computer accidentally practices the exact test it is supposed to take later.
In response, the developers removed that benchmark from their official results. They recalculated everything to provide a fair picture of how well the model actually works. They are now working to ensure this does not happen again.
Comments (0)
No comments yet. Be the first!
More AI news
NewsChatGPT can now access your health data
You can now link your Apple Health and medical records directly to ChatGPT.
NewsSakana AI updates its model router
Sakana AI released version 1.1 of its Fugu Ultra router with performance improvements.
NewsClaude Adds Powerful Features to Voice Mode
You can now use Claude's smartest models to manage your emails and calendar through voice commands.