5 ms·
I have a script that ranks these based on codingindex from Artificial Analysis. All it does is pull a json from their main table page and parses it with the fi
by kristopolous 4mo ago
I have a script that ranks these based on codingindex from Artificial Analysis.
All it does is pull a json from their main table page and parses it with the fields I care about (coding).
There used to be a mailing list associated with it but eh ... there wasn't much interest. I use the script every day though.
Current partial output
score age size name
47.1 58 large Kimi K2.6
47.5 54 large DeepSeek V4 Pro (Reasoning, Max Effort)
47.5 70 - Muse Spark
47.6 132 - Claude Opus 4.6 (Non-reasoning, High Effort)
47.8 205 - Claude Opus 4.5 (Reasoning)
48.1 132 - Claude Opus 4.6 (Adaptive Reasoning, Max Effort)
48.6 55 - GPT-5.5 (Non-reasoning)
48.7 188 - GPT-5.2 (xhigh)
50.1 29 - Qwen3.7 Max
50.7 1 large GLM-5.2 (max)
50.9 120 - Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)
51.5 92 - GPT-5.4 mini (xhigh)
52.1 55 - GPT-5.5 (low)
52.5 62 - Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
53.1 132 - GPT-5.3 Codex (xhigh)
53.1 62 - Claude Opus 4.7 (Non-reasoning, High Effort)
55.5 118 - Gemini 3.1 Pro Preview
56.2 55 - GPT-5.5 (medium)
56.7 20 - Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
57.2 104 - GPT-5.4 (xhigh)
58.5 55 - GPT-5.5 (high)
59.1 55 - GPT-5.5 (xhigh)
62 8 - Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
To see everything, run it like so
$ curl day50.dev/art-analysis.sh | bash
The repo: https://github.com/day50-dev/aa-eval-email https://github.com/day50-dev/aa-eval-email
some key takeaways:
* open models are on about a 4-7 month lag right now depending on how you want to measure it
* if this keeps up, you might see an open-weights model doing claude fable 5 level work before the new year.
if people sign up for the free mailing list (that just does this) I'll go and put it back on ... emails when new model evals drop - it was pretty useful.
- alecco 4mo agoConsider using decrementing score order (best on top)
- kristopolous 4mo agothen I'd have to scroll up over 500 lines after running it every time to see what I care about. But if that's your thing, here you go: https://github.com/day50-dev/aa-eval-email/commit/1853be64610393c5b2a069c610c70d7879a9ebe1 https://github.com/day50-dev/aa-eval-email/commit/1853be6461... add an argument (any argument) and it will be sorted as your specified. It just works as a toggle flipping the order ... so literally any string will do. The original link has been updated accordingly with the new code.
- datadrivenangel 4mo agoHave it print paginated or just top 10?
- kristopolous 4mo agoonly the small ones: $ ./art-analysis.sh | grep small or maybe just the qwen $ ./art-analysis.sh | grep Qwen only the ones in the past 30 days $ ./art-analysis.sh | awk '$2 < 31' I use it in pipes like this.
- spwa4 4mo ago[dead]
- slig 4mo agoThanks for sharing. I'm curious: why didn't you sort with the score descending?
- fridder 4mo agoNot OP but if you run this from the CLI it does make the ordering make a little more sense
- kristopolous 4mo agoBecause it's currently 511 lines. Why would I want to scroll up to see the stuff I care about? Don't you want the relevant stuff to be right there in front of you?
- duckmysick 4mo agoI do and that's why I pipe the output to `head -n 20` or use `LIMIT 20` in SQL. That aside, this is a good script you're running. Thanks.
- tasuki 4mo agoBut maybe you decide you want to see more. It makes perfect sense for a cli tool to output the most interesting piece of info last: then you can decide on the fly whether you want to scroll up or not.
- deleted 4mo ago[deleted]
- slig 4mo agoThank you, that makes a lot of sense.
- snsnbsne 4mo agoBecause programmers can’t figure out how to have a CLI that prints in a normal order, with the newest stuff on top instead of on the bottom. Setup a fresh new large monitor. Open CLI. Run command. Watch output at the bottom of your screen. Keep watching the bottom of your screen for the rest of the day. Sure you can tile windows and it helps but come on. Just have the command/input section in the bottom and the “output” on top. Keep the command bit on the bottom.
- deleted 4mo ago[deleted]
- deleted 4mo ago[deleted]
- deleted 4mo ago[deleted]
- papersail 4mo agoscore age size name 62.0 8 - Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 59.1 55 - GPT-5.5 (xhigh) 58.5 55 - GPT-5.5 (high) 57.2 104 - GPT-5.4 (xhigh) 56.7 20 - Claude Opus 4.8 (Adaptive Reasoning, Max Effort) 56.2 55 - GPT-5.5 (medium) 55.5 118 - Gemini 3.1 Pro Preview 53.1 132 - GPT-5.3 Codex (xhigh) 53.1 62 - Claude Opus 4.7 (Non-reasoning, High Effort) 52.5 62 - Claude Opus 4.7 (Adaptive Reasoning, Max Effort) 52.1 55 - GPT-5.5 (low) 51.5 92 - GPT-5.4 mini (xhigh) 50.9 120 - Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) 50.7 1 large GLM-5.2 (max) 50.1 29 - Qwen3.7 Max 48.7 188 - GPT-5.2 (xhigh) 48.6 55 - GPT-5.5 (Non-reasoning) 48.1 132 - Claude Opus 4.6 (Adaptive Reasoning, Max Effort) 47.8 205 - Claude Opus 4.5 (Reasoning)
- tcp_handshaker 4mo agoShort comments... - GPT 5.5 consistently the best, an opinion who gets me constant downvotes here by the Anthropic Marketeer strike force... - China is going to eat the US lunch on AI - What have European universities and companies been doing? Its like if, on a parallel past/future, Nikola Tesla and Edison would have created flying Cyberpunk machines, while Europeans researchers, would be getting together to request EU funds, for investigation on how to breed faster horses. - If Zuckerberg could be fired, after spending a total of $235 billion on AI and having NOTHING to show for...should he be fired?
- kristopolous 4mo agoThey did muse spark ... it's not garbage. Also what are they building it for? I'd think it's to serve ads better or something like that. Maybe Muse Spark fits facebook's needs perfectly...
- jansan 4mo agoMo Bitar said something like "Meta's LLM is the one you use if you accidentially hit the wrong button in WhatsApp. Its user base is fat-finger phone users."
- bodhi_mind 4mo agoCool project! Side note: Kind of a bad practice imo to ask people to blindly execute bash from an unknown source.
- scrollop 4mo agoWould be interesting to see where gpt 5.5 pro extended is.
- drob518 4mo agoMaybe your script could sort based on score.
- sosodev 4mo agoNote that AA's coding index is only made up of two benchmarks: Terminal-Bench Hard and SciCode. I'm skeptical that it makes a good coding index. It ranks Gemma 4 31B above Deepseek V4 Flash. Having used both of those models for a broad variety of coding tasks I would choose Deepseek every day.
- OkGoDoIt 4mo ago[dead]
- jarjoura 4mo agoSeems legit. My experiments with GLM-5.2 so far have resulted in strange hallucinations in the tiniest of places. Like a wrong variable name. It seems like it's up for the task of complex code, but those little paper-cuts are scary to me. I wouldn't trust this model for anything remotely serious.