3 ms·
There's a table here showing some "Overall" and "Median" score, but no context on what exactly was tested. It appears to be in the ballpark as the latest models
by z2 1y ago
There's a table here showing some "Overall" and "Median" score, but no context on what exactly was tested. It appears to be in the ballpark as the latest models, but with some cost advantages with the downside of being just as slow as the original r1 (likely lots of thinking tokens). https://www.reddit.com/media?url=https%3A%2F%2Fpreview.redd.it%2Fdeepseek-r1-0528-v0-09patvqurk3f1.jpeg%3Fwidth%3D1080%26format%3Dpjpg%26auto%3Dwebp%26s%3D38580df0e04cb8e29921a6a55b24a7f76b5b30ff https://www.reddit.com/media?url=https%3A%2F%2Fpreview.redd....
- xelos 1y agoIt’s appeared on the Livecodebench leaderboard too. Performance on par with O4 Mini - https://livecodebench.github.io/leaderboard.html https://livecodebench.github.io/leaderboard.html