4 ms·
I like to compare models with a similar score on cost per task and output tokens per task since those measure two things I'm interested in: cost efficiency and
by scotttrinh 2mo ago
I like to compare models with a similar score on cost per task and output tokens per task since those measure two things I'm interested in: cost efficiency and token efficiency. Here's how GLM-5.3 compares to other models in a similar score and against GLM-5.2 to save a few clicks for others who care about these metrics:
Model Score Cost / Task Output Tokens / Task
-------------------------------------------------------------------------
GLM-5.3 (max) 59.5 $0.68 41,107
GLM-5.2 (max) 53.0 $0.56 32,200
Claude Opus 5 (high) 61.5 $1.52 21,353
GPT-5.6 Sol (max) 60.9 $1.23 16,879
Grok 4.6 (high) 60.9 $0.84 21,735
Kimi K3 (max) 59.7 $0.84 25,474
GPT-5.6 Sol (xhigh) 59.0 $0.87 11,098
Claude Opus 5 (medium) 58.6 $0.98 12,459
Qwen3.8 Max 58.1 $1.13 38,287
Qwen3.8 2.4T A95B 57.7 $0.95 32,472
Claude Opus 4.8 (max) 57.3 $1.65 33,557
GPT-5.6 Sol (high) 57.3 $0.52 7,545
Muse Spark 1.2 (xhigh) 56.8 $0.40 30,430
GPT-5.6 Terra (max) 56.6 $0.51 20,838
GPT-5.5 (xhigh) 56.3 $0.69 16,893
Gemini 3.7 Flash (high) 56.0 $0.40 36,847
Edited for accuracy and more models.
- sourcecodeplz 2mo agoMuse Spark has a nice balance. not to mentions the Contribs version is old deepseek flash prices.
- sscaryterry 2mo agoI found the sweetspot here: GPT-5.6 Sol (high) 57.3 $0.52 7,545 (Edit: TLDR; It gets on with it, makes the same mistakes you would, without overthinking and overengineering, most of the time)
- glub 2mo agoTested muse spark 1.2 because it was rated so high on design arena, and I've missed a model that can do nice UI in the hands of an operator with no UI skills. It produced worse UI mockups than GPT and GPT models are already the bottom of the barrel here. The only model that performed well was Kimi K3 - insanely good, but expensive. It's hard to trust benchmarks these days.
- samtheprogram 2mo agoIf you just want it to generate UI out of nothing, the benchmarks aren't really for that. If you want to generate a UI based on specific user input of some kind, then they are. I'd suggest using one model for UI and another model for tacking onto that UI. LLMs are great at pattern matching, and benchmarks don't really capture one-shotting desirable UI. That said, benchmaxxing is a thing and your experience with models is a thing. Benchmarks are fuzzy and should be taken with a grain of salt.
- teravor 2mo agothese $/task figures aren't very useful in my experience. it doesn't tell you how well it did the task. generally I choose models by their intelligence and then personal preference from direct experience.
- dudeinhawaii 2mo agoIt would make reading and comparing a bit easier if the data was sorted by a dimension.
- ac29 2mo agoCost per task: Model Score Cost / Task Output Tokens / Task ------------------------------------------------------------------------- Muse Spark 1.2 (xhigh) 56.8 $0.40 30,430 Gemini 3.7 Flash (high) 56.0 $0.40 36,847 GPT-5.6 Terra (max) 56.6 $0.51 20,838 GPT-5.6 Sol (high) 57.3 $0.52 7,545 GLM-5.2 (max) 53.0 $0.56 32,200 GLM-5.3 (max) 59.5 $0.68 41,107 GPT-5.5 (xhigh) 56.3 $0.69 16,893 Grok 4.6 (high) 60.9 $0.84 21,735 Kimi K3 (max) 59.7 $0.84 25,474 GPT-5.6 Sol (xhigh) 59.0 $0.87 11,098 Qwen3.8 2.4T A95B 57.7 $0.95 32,472 Claude Opus 5 (medium) 58.6 $0.98 12,459 Qwen3.8 Max 58.1 $1.13 38,287 GPT-5.6 Sol (max) 60.9 $1.23 16,879 Claude Opus 5 (high) 61.5 $1.52 21,353 Claude Opus 4.8 (max) 57.3 $1.65 33,557 Benchmark score: Model Score Cost / Task Output Tokens / Task ------------------------------------------------------------------------- Claude Opus 5 (high) 61.5 $1.52 21,353 GPT-5.6 Sol (max) 60.9 $1.23 16,879 Grok 4.6 (high) 60.9 $0.84 21,735 Kimi K3 (max) 59.7 $0.84 25,474 GLM-5.3 (max) 59.5 $0.68 41,107 GPT-5.6 Sol (xhigh) 59.0 $0.87 11,098 Claude Opus 5 (medium) 58.6 $0.98 12,459 Qwen3.8 Max 58.1 $1.13 38,287 Qwen3.8 2.4T A95B 57.7 $0.95 32,472 Claude Opus 4.8 (max) 57.3 $1.65 33,557 GPT-5.6 Sol (high) 57.3 $0.52 7,545 Muse Spark 1.2 (xhigh) 56.8 $0.40 30,430 GPT-5.6 Terra (max) 56.6 $0.51 20,838 GPT-5.5 (xhigh) 56.3 $0.69 16,893 Gemini 3.7 Flash (high) 56.0 $0.40 36,847 GLM-5.2 (max) 53.0 $0.56 32,200
- tw1984 2mo agothis is not very useful. for over 1 billion real world users living in China, they don't have the option of paying $1.52 per task to use Opus 5, they are banned doing that due to US politics.
- Barbing 2mo agoWhat’s the reseller situation?