4 ms·
GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
- joehabeebs 3mo agoInteresting tests being done but I can't help but think it limits testing innovation in some way given that the requested apps are essentially all clones of others
- christophilus 3mo agoI hear this take a lot, but every app I’ve ever built was like 80% similar to every other app out there. The unique/ creative part of an app is not the bulk of it, and LLMs have been pretty good at helping me explore the 20%, too.
- billyp-rva 3mo agoCalculator / Rubik's cube / game of life apps should be very close to 100% identical, right? I don't see the point of asking an AI for one of these when there are dozens (hundreds?) of repos that all have exactly what you want.
- sgk284 3mo agoSimilarly, we updated our model arena (52 apps each built by 26 models) to have GPT 5.6 Sol, Terra, and Luna today: https://arena.logic.inc/ https://arena.logic.inc/ It's really interesting to see the Sol/Terra/Luna apps side-by-side. I need to add these stats somewhere in the UI, but one interesting take away: Terra took 1/2 as much wall-clock time as Sol, but Luna took more wall-clock time than Sol (by about 23%). It's still much much cheaper, but it seems like Terra is likely a more optimal time/cost balance for most use cases. The Terra quality is usually nearly as good as Sol, but much faster and cheaper. I do appreciate Sol's design sensibilities (see, for example, the audio sequencer). It's the first model in a while that is clearly distinct on that front. They'd all converged to very similar visuals for a while.
- ianm218 3mo agoThis does seem to validate the critique that models like GLM are benchmaxxed and not as close to the frontier as you’d think based on their numbers.
- ttoinou 3mo ago"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.
- adammarples 3mo agoBecause serious science is hard and valuable for its rigour, and shouldn't be compared with just poking at data to see what happens
- chris_money202 3mo agoDon't be fooled, there is politics, opinions, and less rigor in science as well.
- adammarples 3mo agoThe extent to which their are, is the extent to which that is not science
- ttoinou 3mo agoNo True Scotsman
- sfn42 3mo agoQualitative science is science.
- minimaxir 3mo agoIt's a preemptive defense against methodology cynicism seen often on sites including but not limited to Hacker News. I've been guilty of including such defenses myself over the years because I've gotten annoyed with receiving such cynicism. Look at the top comment on their previous HN submission: https://news.ycombinator.com/item?id=48839886 https://news.ycombinator.com/item?id=48839886
- kibae 3mo agoThe cost seems to be using the wrong symbol: ¢ vs $
- delichon 3mo agoNope, they're that cheap. E.g. Grok 4.5 is $.02 to $.06 per million tokens. A 400 token reply costs ~.002¢ https://www.tryai.dev/models/grok-4.5 https://www.tryai.dev/models/grok-4.5 Update: kibae above and below is correct and I'm not. They have fixed their blog post.
- il 3mo agoGrok 4.5 is $2/$6 there's no model anywhere close to that cheap
- delichon 3mo agoThe numbers come from the tryai.dev link: How much does Grok 4.5 cost on TryAI? Grok 4.5 is Input: $0.02 / 1M tokens, Output: $0.06 / 1M tokens. There is no subscription — you pay only for what you use.
- ricardobeat 3mo agoSlop pricing pages? https://www.tryai.dev/models/claude-fable-5 https://www.tryai.dev/models/claude-fable-5 says Fable costs $0.1/$0.5. Can't wait to use it at those prices! (edit: these have been fixed shortly after)
- kibae 3mo agoThose numbers are incorrect unless they got a deal that's 100x cheaper than API pricing which is unlikely. They updated the original post with the correct costs.
- smusamashah 3mo ago"One honest caveat", "no glitches, no color changes" good tests and I read it to the end but I wish it was written by a human.
- deleted 3mo ago[deleted]
- ranyume 3mo agoYou are absolutely right!
- deleted 3mo ago[deleted]
- BatFastard 3mo agoI hear "Honestly" more often from Anthropic than I ever do from all humans.
- Tadpole9181 3mo agoAnd it uses the word in weird ways I have never experienced. I'm genuinely curious what's in the training set that caused this tic.
- ValentineC 3mo agoClaude training should learn that sometimes people use the word "honestly" because they'd otherwise be lying most of the time.
- TacticalCoder 3mo agoCan't we just take that new language that llmish is and feed it to a transformer of sort that'd get rid of those infuriating sentences? Nothing hard. Everybody wins.
- CharlesW 3mo ago> We generated a big pile of artifacts, we are publishing all of them, and you can form your own opinion. My opinion is that spamming HN with two gimmicky "one-shot prompting shootout" marketing pieces in two days does not build confidence about either your technical or marketing expertise.
- nomel 3mo agoSay you were interviewing a human, to see how capable they were. You are allowed to give them take home work. What kind of questions would you ask, or tasks would you give, to try to get a measure of their competence? If you gave them a task, would you iterate with them on the design, or would you see what they could produce on their own, without input? Measuring "intelligence" is hard, but giving an "intelligent" entity tasks, and seeing what comes out, and then comparing the output with others, seems like a very reasonable, relative, way to do it.
- CharlesW 3mo ago[deleted because pointless]
- nomel 3mo ago> Even if this was a good idea when applied to humans (it's not) I'm not sure I understand. What's not a good idea? I'm asking you how you would do it, with some possible examples. Or, are you saying it's a bad idea to try to measure how competent someone is before hiring them? > LLMs aren't humans, Not sure how this is relevant. My question was how to measure competence and "intelligence" for a task an entity, intelligent enough to do that task, will do. LLMs are not humans, but are usually used to complete tasks humans want completed that would usually be done by humans. That's where the most token spend is for them. Since that's what people are using them for, it seems reasonable to try to measure competence in those tasks.
- thebigspacefuck 3mo ago(LM)Arena is basically this. IMO it’s the best benchmark that avoids benchmaxxing Agent: https://arena.ai/leaderboard/agent https://arena.ai/leaderboard/agent Web dev: https://arena.ai/leaderboard/code/webdev https://arena.ai/leaderboard/code/webdev Currently Fable and 5.6 are neck and neck on web dev which is basically the same finding as this.
- small_model 3mo agoDoesn't have Grok 4.5 listed yet, wonder why 5.6 is, it was released later?
- Chu4eeno 3mo agoThere's a ton of arenamaxxing going on (especially from facebook), though I don't disagree that it's one of the better actual benchmarks. Always fun to ask them to recreate classic demoscene effects (sadly they're still pretty bad at generating music, though at least claude seems to create decent synths). I keep trying to get them to recreate the fluid+particle stuff from Agenda Circling Forth etc., but even giving them the blog posts describing the implementation (and screenshots) they're still pretty bad.
- tedsanders 3mo agoArena can definitely be benchmaxxed a bit, if you try. The distribution of prompts there is very different than usage by regular coders. E.g., lots of requests for one-shot games from scratch. So if you fine-tuned your model to be great at making fun one-shot games from underspecified prompts, your coding model might look better than it is (on general tasks, at least). I work at OpenAI, and am happy to say we don't try to juice our scores here, as doing so would be counterproductive and make Arena a worse signal for everyone.
- rbehrends 3mo agoMy concern with most of these visual benchmarks, popular as they are, is that they are likely more indicative of knowledge (i.e. how comprehensive the training data is and how well it can be retrieved from the model) than of reasoning ability. I don't see in particular how a model would construct a CoT that mapped somehow to a representation of the cube geometry and its animations in latent space without a large chunk of that being pre-existing information.
- nomel 3mo ago> without a large chunk of that being pre-existing information. Is there any evidence that novel reasoning is present in LLM? I've never been able to make that work, and I believe Apple's paper some time ago was good evidence that it doesn't exist. In my experience, sparse latent spaces result in a complete, comical, failure in reasoning.
- drivebyhooting 3mo agoSee the new mathematical proof published by OpenAI. I’m not very valiant to verify its veracity. But even if the math is merely derivative it merits mention.
- nomel 3mo agoTrue, but that's an unknown internal model, without details of the architecture. We'll have to see if the LLM model, itself, was responsible for the "novel" bits, or if it was stuffs bolted onto the LLM that made it possible. I suppose "LLM" is maybe no longer sufficient to describe the systems that LLM are being integrated into, so maybe my point is pedantic/semantic.
- Chu4eeno 3mo agoYeah, Anthropic likely is gaining an edge in tests like this from the data they got from Canva.
- throw310822 3mo ago"Elon and Bezos watch a Blue Origin landing" svgs are super cute, and incredibly like children's drawings. They also nail Bezos' features pretty well.
- sangupta 3mo agoSign-in via Google is broken - it redirects back to localhost from Supabase :)
- NBJack 3mo agoGiven this appears to completely exclude Google's models, I'm not surprised. Even Muse is in there. I guess they aren't fans.
- CompoundEyes 3mo agoIt’s interesting how all the model names and versions are like SKUS taking up space on a display shelf. I look forward to whatever Sagittarius A* does!
- paxys 3mo ago> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and GLM are the slowpokes You put in a lot of good work, and kudos for that, but man, reading paragraphs like these just puts me off of the entire piece. Like…how hard would it have been really to type these two sentences by hand, in your own natural voice?
- jakubmazanec 3mo ago> how hard would it have been really to type these two sentences by hand, in your own natural voice On the other hand, do we have to complain about every seemingly AI written text?
- fluidcruft 3mo agoIt is just horrible writing style. It doesn't particularly matter that AI wrote it.
- jakubmazanec 3mo agoDon't read it then. Why waste enrgy on complaining?
- fluidcruft 3mo agoFor the same reasons you replied to my comment
- paxys 3mo agoYes? AI generated text is explicitly disallowed on HN, so it’s not crazy to expect that a similar standard be used for linked content.
- dinkleberg 3mo agoIs this how I learn that Bezos now has a beard? Interesting that it is a detail that all of the models chose to include (unless that was in the prompt and just not put in the post).
- ricardobeat 3mo agoObviously AI-written, but I'm confused with the results: Muse Spark has the best Rubik's cube by far, the only one properly animating, yet it gets a 2/5 (edit: seems to be an issue with inline videos)
- rsstack 3mo ago2/5 isn't quality, it's consistency as written there. The full links are at the bottom. Most of Spark's attempts are failures: https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/rubiks-cube/muse-spark-1.1/attempt-4.html https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/... https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/rubiks-cube/muse-spark-1.1/attempt-2.html https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/... https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/rubiks-cube/muse-spark-1.1/attempt-5.html https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/...
- ricardobeat 3mo agoAh, I missed that, and didn't click through the links. Most of the videos are not showing any animation for me, only Opus / Qwen / Muse, so Grok's attempt looked broken.
- orliesaurus 3mo agoReally nice breakdown, surprised by the results - especially the fact that OSS models were so behind on most task... (lol at the SVG of the moon without any sign of life by GLM-5.2)
- orliesaurus 3mo agoMissing the exact prompts - would love to replicate...but also curious how you prompted these: they could be a big reason why some models failed completely at rendering SVGs (ie. GLM 5.2)
- platinumrad 3mo agoMaybe I'm a control freak, but asking agents to one-shot random apps is nothing like how I actually use AI in software engineering.
- bhu8 3mo agoAbsolutely yes, but that's how you become twitter/X famous
- fragmede 3mo agoIt's not, but it's trying to bring any level of objective measure in this realm, vs just going off of vibes.
- thsbrown 3mo agoMan I do ponder this all the time.
- mikeocool 3mo agoYeah, the models have all been really good at generating greenfield apps for a really long time (in the scope of LLM time). I suppose it’s interesting to see how they make better greenfield apps. But I am much more interested in how they solve hard problems in existing gnarly codebases.
- ValentineC 3mo agoOne-shot benchmarks are great for me as a solo creator, since they slightly correlate to whether the better frontier models (Opus and Fable for me) make better decisions about things I didn't spec, or whether they'll give me better suggestions right off the bat.
- anuramat 3mo agoI imagine one could one-shot a basic app and then feed feature requests one by one, sounds like an obvious way to benchmark architecture/maintainability
- anuramat 3mo agojust found a decent looking benchmark for iterative development: https://swe-milestone.com/ https://swe-milestone.com/ surprised it isn't a bigger thing, eg artificial analysis doesn't report anything like that still doesn't measure the human-agent interaction part, but that's pure vibes atp
- master_crab 3mo agoA lot of these are visual-heavy tests that often require first person sight to confirm results. Considering GLM isn’t multimodal, that might explain why it did better on the calculator question and not much else.
- esafak 3mo agoCould you make the tables sortable?
- dang 3mo agoRecent and related: We made Grok 4.5, GPT-5.5, and Claude build the same apps - https://news.ycombinator.com/item?id=48838772 https://news.ycombinator.com/item?id=48838772 - July 2026 (92 comments)
- konart 3mo agoNot sure what prompt was used for GLM 5.2 but here is mine: > Draw a horse riding an astronaut in svg https://www.svgviewer.dev/s/if4gi3e7 https://www.svgviewer.dev/s/if4gi3e7
- losvedir 3mo agoI think there's approximately zero value in seeing how a model can turn 100 tokens into a 100k. What workflow is that? It's not useful in the real world. I want to know how well it can follow instructions, manage various potentially competing desires in the context, and so on. It's much more interesting how it can turn 100k tokens (e.g. a codebase and lots of tool calls) into 100 tokens.
- didip 3mo agoI actually like this methodology of testing AI much better than all the other benchmark tests. Real world is messy, other benchmarks are clearly gameable by the Chinese open models. Great job! And I don’t care about the tone of the article, it’s readable just fine.
- deleted 3mo ago[deleted]
- Scroll_Swe 3mo agoGrok redeemed? But I wonder how long they will keep it cheap
- canada_dry 3mo agoI gave up on Grok. It constantly ignores explicit instructions (e.g. do NOT remove existing comments) and it's not nearly as intuitive in knowing the right questions to ask as gpt, claude or gemini in my experience from using all of them.
- spwa4 3mo agoI wish in benchmarks like these people would throw in one or two games that requite creativity. How about "make count binface versus the British reformist space mutants. He has a spaceship, and his head is a trashcan. Make it nice."