5 ms·
Launch HN: Design Arena (YC S25) – Head-to-head AI benchmark for aesthetics
Hi HN, I’m Grace from Design Arena (https://www.designarena.ai/ https://www.designarena.ai/) - we’re building a crowdsourced benchmark for AI-generated visuals (websites, images, video, and more). We put AI models and builder tools in head-to-head comparisons that get voted on by real users from around the world. Think “Hot or Not” for the AI era :)
(Btw, when we say real users we mean real users, so you may get a captcha on the site. Sorry, but we have to use every bot protection available! We only want human ratings, for obvious reasons.)
Here’s a demo video: https://www.youtube.com/watch?v=vPyEQnuVgeI https://www.youtube.com/watch?v=vPyEQnuVgeI
We didn’t set out to build this - we were actually working on an AI game engine. But we found that models sucked at look-and-feel. Even when the output code was usually functional, most visual aspects lacked the soul that makes great graphics feel alive.
So we built a this-or-that game, just for ourselves, to figure out which generated outputs had the best graphics. To our surprise, that turned out to be more exciting than the original idea—it turns out this is a widespread problem! We did a Show HN a month ago (https://news.ycombinator.com/item?id=44542578 https://news.ycombinator.com/item?id=44542578) and that was partly what convinced us to make this benchmark thing our actual product.
State-of-the-art models might be winning IMO gold, but they are still putting white text on a white background. There needs to be some measurement of what’s good and what isn’t (yes, there is such a thing as good design!), and it sure isn’t going to come from LLMs.
We come from engineering backgrounds (Apple and Nvidia) with a love for design; we know when we like or dislike something, even when we can’t say why. This-or-that / hot-or-not games are made for domains like this: Design Arena’s goal is to make everything stupidly simple so humans can just do the easy part: like-vs.-dislike. Which also turns out to be the valuable part, because what’s easiest for humans is actually the part that the AIs can’t currently do.
Since our Show HN, we’ve extended our initial set of ~25 LLM models to 54 LLM models, 12 image models, 4 video models, 22 audio models, and 22 vibe-coding tools (like Lovable, Bolt, v0, Firebase Studio, and more). In this last category, we’ve been surprised to find that agentic tools that were not specifically marketed as vibe-coders like Devin performed exceedingly well in the builder category, outperforming dedicated builder tools like Lovable, v0, and Bolt.
Our users are mostly devs who want to spin up a frontend, or designers who want to spin up design variants faster. In both cases, Design Arena provides a quick way to find out which options are better than others. Dev-or-designer needs to make the final calls, because there’s no substitute for good judgment. But this type of formatting can really help.
We plan to make money by offering version testing as a service to companies that need to quantify improvements in their product between builds.
This is the first time we’ve ever worked on something like this! We’d love to learn from you all and look forward to your feedback.
- transformi 1y agoCool - do you train model that will be the proxy from the votes of persons?
- grace77 1y agowe're not training models or proxying human votes with models
- deleted 1y ago[deleted]
- ryhanshannon 1y agoIs this an area that is not yet covered by other user rating benchmark sites like LLMarena?
- grace77 1y agoyes! LMArena recently started pushing "webdev" arena, but there was no explicit emphasis on design or aesthetics, just web-based content
- KaoruAoiShiho 1y agoCurious if you guys got into YC for this idea or something else?
- neonate 1y agoPost says they were making an AI game engine, so that's probably what they got in with.
- j_da 1y agoWe started out building a platform to one-shot games (single-player and multi-player), but realized that the model you used under the hood really made a difference in functionality and graphics. We started out building the benchmark as an internal tool for ourselves to see which model was the best, but found that benchmarking models on visual "taste" was something that people were generally interested in currently.