6 ms·
Choosing an AI model: one prompt, 11 models, different results
- isqueiros 2mo ago> Build a one-page site for a neighbourhood coffee shop: opening hours, the address, a short menu and a photo. Nothing on it changes unless I edit it myself. If that's the entire prompt, it's quite depressing how much alike these all look. I appreciate some of the details from the Opus 5 version, but I can't help but strongly feel the AI vibes emanating from that design.
- hombre_fatal 2mo agoWithout more knowledge about the technical aspect, it could be a good thing that they're all so similar. If you tell humans to go from their kitchen to their bedroom, they all stand up and walk the same way. Nobody decides to crab walk. Maybe you'd have to cripple the model in some way to do that. On the other hand, if you want something different with LLMs, all it takes is a few more words of creative flair in the prompt.
- skinfaxi 2mo agoThis is an interesting point. But even then everyone has their own style. Some might walk with a bit of a swagger, some with a limp, some might have to get in a wheelchair to go over. Would model temperature be another knob to turn beyond a more creative prompt?
- mym1990 2mo agoThe models are giving you very average, middle of the road responses. The average human gait doesn't have swagger, limps, or extensive accessibility needs. If you want the swagger, you need to steer it that way.
- hombre_fatal 2mo ago> with a limp > wheelchair You really took "cripple the model" to heart!
- skinfaxi 2mo agoThat wasn't my intention and I sincerely did not mean any offense.
- skeledrew 2mo agoI don't think they took offense; more likely what they meant was "limping" and "wheelchair" aren't considered normal ways to get around, because there's an abnormal condition that led to said ways of movement. Similarly LLMs with similar capabilities operating in a normal way should produce relatively similar output.
- tantalor 2mo agoIt's also kind of pointless. Why does a coffee shop need a website anyway? Nobody is saying, "man I would love to go to this coffee shop but I can't find their website". If all you want is "opening hours, the address, a short menu and a photo" there are easier and cheaper ways to do that.
- miyoji 2mo ago> Nobody is saying, "man I would love to go to this coffee shop but I can't find their website". Me, I'm saying that, and I've skipped going to coffee shops and restaurants because they don't have a website, just a fucking Facebook page. I don't use Meta products and can't see their page if I'm not logged into an account I don't have, so I do what the business owner intended: I go fuck myself and get coffee somewhere else.
- altmanaltman 2mo agoBut... why do you want to look up a coffee shop online before going there? Honestly have never heard anyone say this in my life before
- skinfaxi 2mo agoTo see the hours, get a sense of the menu. There are many matcha shops by me for instance and my wife likes to check the specials before choosing which one to go to.
- altmanaltman 2mo agoDoesn't most of this show up on google anyway?
- mym1990 2mo agoIt sounds like you know everything about the internet and people's preferences, why ask?
- andy99 2mo agoIt’s always been interesting to see how similar output is across ostensibly very different models. I remember testing short story writing in the early days and having the models all choose the same niche topic across e.g GPT, Llama, Phi, Claude, etc
- jamiedumont 2mo agoI trialed Claude for a sorely needed redesign on my website. I was quite impressed with the results, but decided to hang fire on deploying it. Cue my surprise when I just found the same design applied to a coffee shop! The similarities are beyond coincidence, to the point I'll be scrapping Claude's version of the redesign.
- regexorcist 2mo agoWhen I set up my local AI server I tried different harnesses and one prompt was to build a site for a consultancy in my domain. Not long after an actual agency I know had a website redesign and it was the exact same thing I got from my local model and the Hermes agent with the web design skill.
- Schlagbohrer 2mo agoA user, even without much design background, could ask the AI in a follow up prompt to change a few key parameters like overall color and differentiate it more. Even taking a sample image from a design magazine or blog and saying "incorporate some of the design elements here" would probably work great to diverge it.
- thatmf 2mo agoI found it pretty interesting. To me, the GPT Sol and Luna ones are the nicest, which is not what I was expecting. Font choice goes a long way. The Opus ones are acceptable, but still disappointing. The rest are hideous.
- giarc 2mo agoI was thinking the same thing, but I suspect there was something more to it. Check out this sentence from the page. >Our default skills also include some UI design guidance, mainly to avoid known gotchas (e.g., the now-dreaded purple AI slop) and get the model to reason about the visual identity appropriate for the user’s ask. But beyond that, each model is free to go build what it thinks we’ll want.
- fasterik 2mo agoI think this is exactly what we should expect. LLMs are engineered to produce the most likely output given the training data. When you give it a generic design task, you're going to get a distillation of design cliches and tropes from the past decade.
- forgotTheLast 2mo agoDon't blame LLMs if the zeitgeist for coffee shop websites is sepia tones and a Papyrus-like font. My university's Principles of WebDev course 10 years ago had a similar assignment and the results all ended up looking like that too.
- NichoPaolucci 2mo agoReminds me of the Bootstrap era. I can't count the amount of websites with a centered header navigation, slightly rounded accent color buttons, hero section that was mostly text, and some "fun quirk" in the background, either geometric shapes or squiggles or something. I think websites have always looked mostly alike. It's sorta always been a thing. Reminds me of "Corporate Memphis" (https://en.wikipedia.org/wiki/Corporate_Memphis https://en.wikipedia.org/wiki/Corporate_Memphis) It's alright though, because some people are OK with middle of the road (Wordpress, Boostrap, Squarespace templates, now AI). And others are willing to either pay a developer to get involved or put in the extra effort to differentiate themselves.
- kifler 2mo agoI always enjoy these comparisons between models, especially when they demonstrate the actual costs in addition to the outputs.
- arjie 2mo agoBuilding ad-hoc evals is trivial these days. And you can place an LLM judge in front to disambiguate between output. e.g. I have Frigate gating a video feed so that an LLM can watch our home cameras and the labeling task takes a few minutes and then the evals run at fairly low cost https://wiki.roshangeorge.dev/w/images/4/4e/Screenshot_-_Eval_Labeler_Home_Cameras.png https://wiki.roshangeorge.dev/w/images/4/4e/Screenshot_-_Eva... This means that generic benchmarks and evals are sort of passé. There's no need to have them write "make a coffee shop page" or whatever. Instead, simply focus on the real problem you have and evaluate whether they solve that problem. You want to move along the price-performance frontier alone and the SOTA models are decidedly at the top of the performance curve but very expensive so they serve very well to produce the gold standard and to judge. The change from prior to today is that benchmaxxing and per-token pricing allowing for evaluation point to the same direction: do not proxy results. Instead deploy the highest end model you have as a judge for others until you have statistical confidence in discrimination and move along the task-specific frontier.
- epolanski 2mo ago> Building ad-hoc evals is trivial these days. I highly doubt that you have solved it. Writing proper evals and benchmarks for real-world scenarios is far from trivial. Benchmarking an agent essentially means freezing, at the very minimum: - the model - the model's configuration (e.g. effort, permissions, provider) - the dataset (e.g. a git repository at a specific sha) - the code running the agent itself (you can build your own harness, trivial, but you still need to ship it as a single executable, froze in time. benchmarking against closed source runtime like claude code is quite useless, they change too frequently and in ways you cannot directly inspect). - the tools at agent's disposal. Even a slightly different implementation of tool X (e.g. grep or readfile or sed) has an impact. In general this implies also freezing a very specific container image. In my personal benchmarks I provide a specific list of tools that come with the executable, there's no possibility of interacting with the outside world besides the provided apis, the agent bundles its own tools. And even then: there's significant noise coming from the LLM providers themselves which noticeably change the models behaviour, I don't know whether this is because they optimize some settings or change the inference over time, etc. And, last but not least, the output of LLMs is non deterministic. Also, the LLM as judge presents essentially the same non-deterministic problems, has to be benchmarked itself thoroughly, and writing quality rubrics or "golden answers/outputs" is just difficult. One of the metrics I consistently try to emphasize is to avoid the "shotgun vomit dump" of information. So answers that get right to the point in plain terms avoiding dumps of information filled of jargon on top of the user are rated differently. In short: its far from trivial to benchmark models on real-world agentic work taken from your personal or professional projects. And even creating the test cases themselves is hard. No: you cannot take the output of some "sota" and use it as gold standard. This is a very crap approach. It's the sloppiest solution to the problem, in the very sense of slop: plausible, average, lacking any creativity or out of the box thinking, the things that make the real difference in complex software development. The very point of creating these benchmarks is to find which configuration/model/tools/harness (skills/mcps/documentation/agents.md, etc) works better. And it only works if you create these benchmarks yourself from genuinely difficult non-trivial work and find a solution that is better by most metrics implementation-wise, albeit you could settle on the implementation solving a series of cases and edge cases.
- plumb_samji 2mo agoInteresting exploratory comparison, but I be cautious about treating it as a model benchmark With only three runs per model, the results are highly sensitive to randomness
- neom 2mo agoI'd be curious to see Terra xhigh vs Sol low, only in that the visual languages are actually kinda different, that Terra work was a lot less "llm" feeling than a lot of the other designs, wonder if pumping up the effort would result in a more in-depth design but within that style.FWIW I enjoyed reading this way more than any usual benchmark posts we see.
- jwr 2mo agoI've been doing a lot of benchmarking for a long time now with a number of local models for the purposes of spam filtering. The major observation is that there is a lot of variance in model performance. This should not be surprising, as these are probabilistic machines based on random numbers, so your performance will vary from run to run. But this also means that any sort of evaluation of benchmark with a sample size of 1 is essentially worthless for the purposes of model comparison. In my benchmarks, I started insisting on having at least 5 runs. This also makes me very suspicious of some of the posted benchmarks that compare models. If the results don't say how many benchmark runs were performed and what the variance was, there is really no basis for comparison. You are just guessing.
- toddmorey 2mo agoYeah the difference in token usage across difference models they found in the article was broader than I expected, but I was sort of floored by the variance in token usage for the same prompt (run multiple times) with the same model.
- touristtam 2mo agoHave you published any of your results? I am quite curious.
- Systemerror7A69 2mo agoAm I wrong or are these evaluations, while interesting, not really meaningful for anyone doing serious development work? I'm asking because I personally only use AI with specific and detailed instructions, building my projects piece-by-piece. I mostly don't look at the low level code and some of it I don't understand as much as I'd like, but I very much give much more technical instructions than a simple, two sentence prompt. So, it seems to me this "oneshot from a simple prompt" eval is fairly meaningless when it comes to model evaluation itself, as it is in no way representative of real world application. This would be more something for "vibe coders", people with little to no programming background wanting a website?
- 217 2mo agoI'm repeatedly noticing that people working at big ai and tech companies are surprisingly not that... good... at using ai? It's like theyre doing a plausible thing to get something done and calling it a day
- maccard 2mo agoCan you share some posts of good examples of prompts and comparisons?
- toddmorey 2mo agoAs someone who's done web development for 20+ years, I find the model personalities pretty dang fascinating, especially how they develop (and evolve) design sensibilities. I'm interested to know how the classic AI "purple preference" emerged (organically?) and if the beige wave came out of specific training efforts to combat it? To your point on development work (the code itself), I was talking to some friends on the Google Chrome team about any research understanding the model's preferences around framework ergonomics and abilities to properly implement core web standards for given tasks. I think that would be super fascinating.
- runtime_terror 2mo agoI'd assume it comes from Tailwind boilerplate/template sites as it seemed almost all of them were purple at the time
- horsawlarway 2mo agoApproaching this from the perspective of a potential customer and not a designer - I find that I actually like the smaller/open model output quite a bit more. Kimi, GLM, and Deepseek all absolutely run away with the "Can I quickly read the menu and find the address" challenge. Most of the rest of the pages are stylistic, but hard to parse. If I were in a car on a mobile phone trying to find the address of the place to meet a friend for coffee... I don't want a bunch of fluff and stylistic design that makes it hard to parse the information on the site.
- jannishan 2mo agoI think we should establish a professional evaluation organization for models; otherwise, it will be difficult for informal evaluations to form standards.
- pedrosbmartins 2mo agoPretty interesting how DeepSeek V4 Flash 0731 has such disparate (and cheap!) results. I would never guess they come from the same model and prompt.
- chrisjj 2mo ago> Vector graphics actually require a lot of work from the models How so? Surely they can just steal such generic graphics off existing web sites.
- deleted 2mo ago[deleted]
- sceptic123 2mo agoMy "favourite" site was <https://6a6fa3376288679d094a8437--ar-testing-coffee-2b15c6a08894.netlify.app/ https://6a6fa3376288679d094a8437--ar-testing-coffee-2b15c6a0...> with footer image captions that are amazing. Precision Late Art being the best one.
- s4i 2mo agoIn my opinion, this kind of a benchmark doesn't tell much about the models' capabilities on normal software development tasks. It's fun to look at the differences in the output of course, but how often would anyone prompt with very brief instructions, without even hinting the model about caring about any of the details in the outcome nor the implementation? When there are implicit boundaries, negotiable tradeoffs, taste, whatever, in the mix, then the differences in the model capabilities become way more interesting.
- TrustScoreAgent 2mo ago[flagged]
- tealmoonx 2mo ago[dead]
- edgyquant 2mo agoBenchmarks are so difficult with ai because as soon as one gets popular it enters the dataset so the next iteration of the model is trained on the solution. I’m not sure if there’s any potential work around here or anyone doing interesting work but would love to hear about it if so
- Schlagbohrer 2mo agoBut if this effect were really that strong, the models should be getting nearly 100% on the common benchmarks. But for many of the benchmarks, even after being public for more than a year, the new models only get 60-80%.
- throwa356262 2mo agoThis seems to be the easiest way to get on HN front page: Come up with an arbitrary test, let a bunch of LLMs work on it. Make some very subjective judgement about the result...
- deleted 2mo ago[deleted]
- Schlagbohrer 2mo agoFinally some really useful apples-to-apples comparisons rather than endless benchmarkmaxxing and anecdata. I want to see this test run again for the open-weights models that fit into 128GB combined RAM. Like how does Muse-glimmer compare with qwen3.6-35B?
- feor 2mo agoI like the output of the cheaper/older models better, surprisingly, say Gemini 3.1 or DeepSeek V4. They're mostly no frills and just text, and ironically look less AI-generated to me because of the lack of hip slogans and graphics. Definitely closer to what I'd want for my own site, but I don't claim to know what people want from a website for a coffeeshop.
- Damjanski 2mo agothis is fun!
- reindeer2 2mo ago[flagged]
- sorokod 2mo agoSame model over 11 days, one prompt, different results?
- nullzzz 2mo agoThis would be interesting, if I was into building coffee shop sites and todo apps from scratch. For the HN crowd tho, I’d say these are toy examples. No offense!
- andai 2mo ago> Of course, Opus will also perform relentless self-validation of its own work (it does not bill itself on good looks alone). But remember there’s certainly a higher-than-average credit cost attached to that. I haven't tried it but I think this can be replicated with a system prompt. I remember the codex system prompt contains something like, "Do not consider a task complete until you have verified the result." Although I've been running the new GPT models in a custom harness and they do that anyway now, without being prompted. So I think that prompt was for a previous generation.
- sinuhe69 2mo agoHaving skimmed through the article, I assumed that they hadn't tested the design for mobile devices because it wasn't mentioned. That would be a huge mistake. In today's web design, a mobile-first approach is imperative. This is even more important if you want to showcase your café. I used the developer tool in Firefox to see how this design would behave on mobile devices. There are huge differences. Some designs use the screen estate so ineffectively that only the title and a big, boring, generic graphic is shown on a phone. Users have to scroll all the way down to see the content and find what they need. Better designs show the menu, navigation points, and meaningful, aesthetic graphics. Other designs, such as Gemini 3.6, were quite sophisticated but not optimized for traffic and would not load on a 3G connection. However, a simple static website should load instantly on a mobile connection. That said, even a simple web page has many requirements, so expecting a turnkey, ready-made design if the user is not guiding the process is not realistic. Thus, I believe the best choice nowadays is a model with good design skills that understands and adheres to an iterative design process, offering a good initial design as a starting point but also prompting the user to provide guidance and feedback. As the design process runs through many cycles, the initial cost should be modest. But more importantly, the model should understand its own design, be able to explain its choices so it can converse with the user using concrete elements in a accurate language to guide the process. IMO, this is still missing from even the frontier models today.
- whateveracct 2mo agoW: X, Y, Z LLM
- ben8bit 2mo agoSo what we have here is the LLM version of the aesthetic usability effect. So: prettier = better. Take from that what you will - and maybe the AUE is more desirable - but Sol & K3 are genuine game changers for existing codebases. Anthropic currently have an antagonism problem which has worked for them in the past, but not anymore I think, as other models have become as competent.
- 44za12 2mo agoBuilt rightsize exactly for this. https://nehmeailabs.com/right-size https://nehmeailabs.com/right-size
- AndroVertex 2mo ago[flagged]
- AndroVertex 2mo ago[flagged]
- danpalmer 2mo agoThat prompt is unrealistic. Shops already have opening hours, prices, locations, etc. If you leave all those up to the model then there's no correct answer and the models will perform better – the less you constrain them the more they can output the median answer. If you constrain them with all those real requirements, they can get things wrong and we can see where they actually fall short.
- swyx 2mo agonice to see you still on the netlify train todd!!!! warms the hear to see netlify back on HN
- jobuildsstuff 2mo ago[flagged]
- amelius 2mo agoNot a fan of N=1 experiments.
- felixlu2026 2mo ago[dead]
- raffraffraff 2mo agoI'm doing a small local version of this test to pull moods and themes out of song lyrics. I've found that one prompt, 1 model, 11 runs even gives different results. There's no consistency over multiple runs of the same prompt on the same model. Also, for the purpose of music lyric analysis Qwen 4B is laughably bad. Like it's going out of its way to be extremely wrong, misunderstand the prompt. When it does correctly understand what I asked for it ALWAYS tells me that the mood is Angry. Sometimesiit just gives back all the lyrics. Sometimes it claims that it doesn't have the list of moods or the lyrics and tells me I should look them up on the Internet first. All with the same prompt every time. Are models being overtuned for coding tasks?