15 ms·
AI World Clocks
"Every minute, a new clock is rendered by nine different AI models."
- orly01 11mo agoWhat does it mean that each model is allowed 2000 tokens to generate its clock?
- zkmon 11mo agoWhy are Deepseek and Kimi are beating other models by so much margin? Is this to do with their specialization for this task?
- deleted 11mo ago[deleted]
- deleted 1y ago[deleted]
- kfarr 11mo agoAdd some voting and you got yourself an AI World Clock arena! https://artificialanalysis.ai/image/arena https://artificialanalysis.ai/image/arena
- BrandoElFollito 11mo agoThank you very much.... It was a fun game until I got to the prompt Place a baby elephant in the green chair I cannot unsee what I saw and it is 21:30 here so I have an hour or so to eliminate the picture from my mind or I will have nightmares.
- syx 11mo agoI’m very curious about the monthly bill for such a creative project, surely some of these are pre rendered?
- coffeecoders 11mo agoNapkin math: 9 AIs × 43,200 minutes = 388,800 requests/month 388,800 requests × 200 tokens = 77,760,000 tokens/month ≈ 78M tokens Cost varies from 10 cents to $1 per 1M tokens. Using the mid-price, the cost is around $50/month. --- Hopefully, the OP has this endpoint protected - https://clocks.brianmoore.com/api/clocks?time=11:19AM https://clocks.brianmoore.com/api/clocks?time=11:19AM
- whimsicalism 11mo agoi think it is cached on the minute level, responses cannot be that fast
- fouc 11mo agoIt was limited to 2,000 tokens each. I assume it usually hit that. So could be closer to 777M. assuming they didn't just cache it and just start rotating after a day or two..
- ugh123 11mo agoCool, and marginally informative on the current state of things. but kind of a waste of energy given everything is re-done every minute to compare. We'd probably only need a handful of each to see the meaningful differences.
- whoisjuan 11mo agoIt's actually quite fascinating if you watch it for 5 minutes. Some models are overall bad, but others nail it in one minute and butcher it in the next. It's perhaps the best example I have seen of model drift driven by just small, seemingly unimportant changes to the prompt.
- alister 11mo ago> model drift driven by just small, seemingly unimportant changes to the prompt What changes to the prompt are you referring to? According the comment on the site, the prompt is the following: Create HTML/CSS of an analog clock showing ${time}. Include numbers (or numerals) if you wish, and have a CSS animated second hand. Make it responsive and use a white background. Return ONLY the HTML/CSS code with no markdown formatting. The prompt doesn't seem to change.
- sambaumann 11mo agopresumably the time is replaced with the actual current time at each generation. I wonder if they are actually generated every minute or if all 6480 permutations (720 minutes in a day * 9 llms) were generated and just show on a schedule
- whoisjuan 11mo agoThe time given to the model. So the difference between two generations is just somethng trivially different like: "12:35" vs 12:36"
- moffkalast 11mo agoKimi seems the only reliable one which is a bit surprising, and GPT 4o is consistently better than GPT 5 which on the other hand is unfortunately not surprising at all.
- PeterStuer 11mo agoWhy? This is diagonal to how LLM's work, and trivially solved by a minimal hybrid front/sub system.
- em3rgent0rdr 11mo agoTo gauge.
- bayindirh 11mo agoBecause, LLMs are touted to be the silver bullet of silver bullets. Built upon world's knowledge, and with the capacity to call upon updated information with agents, they are ought to rival the top programmers 3 days ago.
- awkwam 11mo agoThey might be touted like that but it seems like you don't understand how they work. The example in the article shows that the prompt is limiting the LLM by giving it access to only 2000 tokens and also saying "ONLY OUTPUT ...". This is like me asking you to solve the same problem but forcing you do de-activate half of your brain + forget any programming experience you have. It's just stupid.
- bayindirh 11mo ago> like you don't understand how they work. I would not make such assumptions. > The example in the article shows that the prompt is limiting the LLM by giving it access to only 2000 tokens and also saying "ONLY OUTPUT ..." The site is pretty simple, method is pretty straightforward. If you believe this is unfair, you can always build one yourself. > It's just stupid. No, it's a great way of testing things within constraints.
- em3rgent0rdr 11mo agoMost look like they were done by a beginner programmer on crack, but every once in a while a correct one appears.
- morkalork 11mo agoI'd say more like a blind programmer in the early stages of dementia. Able to write code, unable to form a mental image of what it would render as and can't see the final result.
- pixl97 11mo agoDeepSeek and Kimi seem to have correct ones most of the time I've looked.
- em3rgent0rdr 11mo agoyes, and sometimes Grok.
- pixl97 11mo agoThe hour hand commonly seems off on Grok.
- BrandoElFollito 11mo agoDeepSeek told me that it cannot generate pictures and suggested code (which is very different)
- shafoshaf 11mo agoIt's interesting how drawing a clock is one of the primary signals for dementia. https://www.verywellhealth.com/the-clock-drawing-test-98619 https://www.verywellhealth.com/the-clock-drawing-test-98619
- BrandoElFollito 11mo agoThis is very interesting, thank you. I could not get to the store because of the cookie banner that does not work (at left on mobile chrome and ff). The Internet Archive page: https://archive.ph/qz4ep https://archive.ph/qz4ep I wonder how this test could be modified for people that have neurological problems - my father's hands shake a lot but I would like to try the test on him (I do not have suspicions, just curious). I passed it :)
- larodi 11mo agowould be gr8t to also see the prompt this was done with
- creade 11mo agoThe ? has "Create HTML/CSS of an analog clock showing ${time}. Include numbers (or numerals) if you wish, and have a CSS animated second hand. Make it responsive and use a white background. Return ONLY the HTML/CSS code with no markdown formatting."
- larodi 11mo agoHmm nothing fancy then, but perhaps with tubing results will vary. I hate prompt discovery (not engineering this thing!), but it actually matters.
- Gormanu 11mo ago[dead]
- bananatron 11mo agogrok's looks like one of those clocks you'd find at a novelty shop
- AlfredBarnes 11mo agoIts cool to see them get it right .....sometimes
- baltimore 11mo agoSince the first (good) image generation models became available, I've been trying to get them to generate an image of a clock with 13 instead of the usual 12 hour divisions. I have not been successful. Usually they will just replace the "12" with a "13" and/or mess up the clock face in some other way. I'd be interested if anyone else is successful. Share how you did it!
- snek_case 11mo agoFrom my experience they quickly fail to understand anything beyond a superficial description of the image you want.
- atorodius 11mo agoThat's less and less true https://minimaxir.com/2025/11/nano-banana-prompts/ https://minimaxir.com/2025/11/nano-banana-prompts/
- dang 11mo agoRelated ongoing thread: Nano Banana can be prompt engineered for nuanced AI image generation - https://news.ycombinator.com/item?id=45917875 https://news.ycombinator.com/item?id=45917875 - Nov 2025 (214 comments)
- Scene_Cast2 11mo agoI've noticed that image models are particularly bad at modifying popular concepts in novel ways (way worse "generalization" than what I observe in language models).
- emp17344 11mo agoMaybe LLMs always fail to generalize outside their data set, and it’s just less noticeable with written language.
- 11mo ago
- abathologist 11mo agoThis is great. If you think that the phenomena of human-like text generation evinces human-like intelligence, then this should be taken to evince that the systems likely have dementia. https://en.wikipedia.org/wiki/Montreal_Cognitive_Assessment https://en.wikipedia.org/wiki/Montreal_Cognitive_Assessment
- AIorNot 11mo agoImagine if I asked you to draw as pixels and operate a clock via html or create a jpeg with a pencil and paper and have it be accurate.. I suspect your handcoded work to be off by an order of magnitutde compared
- jonplackett 11mo agokimi is kicking ass
- busymom0 11mo agoBecause a new clock is generated every minute, looks like simply changing the time by a digit causes the result to be significantly different from the previous iteration.
- shevy-java 11mo agoNow that is actually creative. Granted, it is not a clock - but it could be art. It looks like a Picasso. When he was drunk. And took some LSD.
- kburman 11mo agoThese types of tests are fundamentally flawed. I was able to create perfect clock using gemini 2.5 pro - https://gemini.google.com/share/136f07a0fa78 https://gemini.google.com/share/136f07a0fa78
- sinak 11mo agoHow are they flawed?
- earthnail 11mo agoThe results are not reproducable, as evidenced by parent poster.
- micromacrofoot 11mo agoisn't that kind of the point of non-determinism?
- earthnail 11mo agoNo. Good nondeterministic models reproducibly generate equally desirable output - not identical output, but interchangeable.
- micromacrofoot 11mo agooh I see, thank you for clarifying
- jmdeon 11mo agoAren't they attempting to also display current time though? Your share is a clock starting at midnight/noon. Kimi K2 seems to be the best on each refresh.
- Drew_ 11mo agoThe website is regenerating the clocks every minute. When I opened it, Gemini 2.5 was the only working one. Now, they are all broken. Also, your example is not showing the current time.
- lxe 11mo agoHonestly, I think if you track the performance of each over time, since these get regenerated once in a while, you can then have a very, very useful and cohesive benchmark.
- 1yvino 11mo agoi wonder kwen prompt woud look like hallucination?
- fschuett 11mo agoReminds me of this: https://www.youtube.com/watch?v=OGbhJjXl9Rk https://www.youtube.com/watch?v=OGbhJjXl9Rk
- S0y 11mo agoTo be fair, This is a deceptively hard task.
- bobbylarrybobby 11mo agoWithout AI assistance, this should take ~10–15 minutes for a human. Maybe add 5 minutes if you're not allowed to use d3.
- alexmorley 11mo agoIt's just html/css so no js at all let alone d3.
- postalrat 11mo agoWhats your hourly rate? I'll pay you to make as many as you can in a few hours if you share the video.
- Mashimo 11mo agoI would not even know how to draw a circle with CSS to be honest.
- Bolwin 11mo agoPretty sure css has a sin() fn, that's half your work
- bobbylarrybobby 11mo agoA div with width = height = border-radius (or width = height with border-radius:50%)
- deleted 11mo ago[deleted]
- zkmon 11mo agoWas Claude banned from this Olympics?
- giancarlostoro 11mo agoHaiku is the lightweight Claude model, I'm not sure why they picked the weaker model.
- collimarco 11mo agoIn any case those clocks are all extremely inaccurate, even if AI could build a decent UI (which is not the case). Some months ago I published this site for fun: https://timeutc.com https://timeutc.com There's a lot of code involved to make it precise to the ms, including adjusting based on network delay, frame refresh rate instead of using setTimeout and much more. If you are curious take a look at the source code.
- mstipetic 11mo agoGPT-5 is embarrassing itself. Kimi and DeepSeek are very consistently good. Wild that you can just download these models.
- shubham_zingle 11mo agonot sure about the accuracy though, although shooting in the dark
- awkwam 11mo agoLimiting the model to only use 2000 tokens while also asking it to output ONLY HTML/CSS is just stupid. It's like asking a programmer to perform the same task while removing half their brain and also forget about their programming experience. This is a stupid and meaningless benchmark.
- system2 11mo agoAsk Claude or ChatGPT to write it in Python, and you will see what they are capable of. HTML + CSS has never been the strong suit of any of these models.
- camalouu 11mo agoClaude generates some js/css stuff even when i don't ask for it. I think Claude itself at least believes he is good at this.
- munro 11mo agoAmazing, some people are so enamored with LLMs who use them for soft outcomes, and disagree with me when I say be careful they're not perfect -- this is such a great non technical way to explain the reality I'm seeing when using on hard outcome coding/logic tasks. "Hey this test is failing", LLM deletes test, "FIXED!"
- worldsayshi 11mo agoYeah it seems crazy to use LLM on any task where the output can't be easily verified.
- palmotea 11mo ago> Yeah it seems crazy to use LLM on any task where the output can't be easily verified. I disagree, those tasks are perfect for LLMs, since a bug you can't verify isn't a problem when vibecoding.
- mopsi 11mo ago> "Hey this test is failing", LLM deletes test, "FIXED!" A nice continuation of the tradition of folk stories about supernatural entities like teapots or lamps that grant wishes and take them literally. "And that's why, kids, you should always review your AI-assisted commits."
- derbOac 11mo agoSomething that struck me when I was looking at the clocks is that we know what a clock is supposed to look and act like. What about when we don't know what it's supposed to look like? Lately I've been wrestling with the fact that unlike, say, a generalized linear model fit to data with some inferential theory, we don't have a theory or model for the uncertainty about LLM products. We recognize when it's off about things we know are off, but don't have a way to estimate when it's off other than to check it against reality, which is probably the exception to how it's used rather than the rule.
- ehnto 11mo ago
- novemp 11mo agoOh cool, it's the schizophrenia clock-drawing test but for AI.
- otterley 11mo agoWatching this over the past few minutes, it looks like Kimi K2 generates the best clock face most consistently. I'd never heard of that model before today! Qwen 2.5's clocks, on the other hand, look like they never make it out of the womb.
- bArray 11mo agoIt could be that the prompt is accidentally (or purposefully) more optimised for Kimi K2, or that Kimi K2 is better trained on this particular data. LLM's need "prompt engineers" for a reason to get the most out of a particular model.
- energy123 11mo agoGoes to show the "frontier" is not really one frontier. It's a social/mathematical construct that's useful for a broad comparison, but if you have a niche task, there's no substitute for trying the different models.
- observationist 11mo agoIt's not fair to use prompts tailored to a particular model when doing comparisons like this - one shot results that generalize across a domain demonstrate solid knowledge of the domain. You can use prompting and context hacking to get any particular model to behave pseudo-competently in almost any domain, even the tiny <1B models, for some set of questions. You could include an entire framework and model for rendering clocks and times that allowed all 9 models to perform fairly well. This experiment, however, clearly states the goal with this prompt: `Create HTML/CSS of an analog clock showing ${time}. Include numbers (or numerals) if you wish, and have a CSS animated second hand. Make it responsive and use a white background. Return ONLY the HTML/CSS code with no markdown formatting.` An LLM should be able to interpret that, and should be able to perform a wide range of tasks in that same style - countdown timers, clocks, calendars, floating quote bubble cycling through list of 100 pithy quotations, etc. Individual, clearly defined elements should have complex representations in latent space that correspond to the human understanding of those elements. Tasks and operations and goals should likewise align with our understanding. Qwen 2.5 and some others clearly aren't modeling clocks very well, or maybe the html/css rendering latents are broken. If you pick a semantic axis(like analog clocks), you can run a suite of tests to demonstrate their understanding by using limited one-shot interactions. Reasoning models can adapt on the fly, and are capable of cheating - one shots might have crappy representations for some contexts, but after a lot of repetition and refinement, as long as there's a stable, well represented proxy for quality somewhere in the semantics it understands, it can deconstruct a task to fundamentals and eventually reach high quality output. These type of tests also allow us to identify mode collapses - you can use complex sophisticated prompting to get most image models to produce accurate analog clocks displaying any time, but in the simple one shot tests, the models tend to only be able to produce the time 10:10, and you'll get wild artifacts and distortions if you try to force any other configuration of hands. Image models are so bad at hands that they couldn't even get clock hands right, until recently anyway. Nano banana and some other models are much better at avoiding mode collapses, and can traverse complex and sophisticated compositions smoothly. You want that same sort of semantic generalization in text generating models, so hopefully some of the techniques cross over to other modalities. I keep hoping they'll be able to use SAE or some form of analysis on static weight distributions in order to uncover some sort of structural feature of mode collapse, with a taxonomy of different failure modes and causes, like limited data, or corrupt/poisoned data, and so on. Seems like if you had that, you could deliberately iterate on, correct issues, or generate supporting training material to offset big distortions in a model.
- deleted 11mo ago[deleted]
- earth2mars 11mo agohttps://gemini.google.com/share/00967146a995 https://gemini.google.com/share/00967146a995 works perfectly fine with gemini 2.5 pro
- lanewinfield 11mo agonice. I restrict to 2000 tokens for mine, how many was that?
- esafak 11mo agohow do you do that?
- earth2mars 11mo agoI used exactly the same prompt this site uses. Nothing else.
- agildehaus 11mo agoI'm assuming the "Gemini 2.5" referenced on this site is Flash, not Pro. Pro is insane, and 3.0 is just around the corner.
- lanewinfield 11mo agohi, I made this. thank you for posting. I love clocks and I love finding the edges of what any given technology is capable of. I've watched this for many hours and Kimi frequently gets the most accurate clock but also the least variation and is most boring. Qwen is often times the most insane and makes me laugh. Which one is "better?"
- anigbrowl 11mo agoI really like this. The broken ones are sometimes just failures, but sometimes provide intriguing new design ideas.
- jdiff 11mo agoThis same principle is why my favorite image generation model is the earlier models from 2019-2020 where they could only reliably generate soup. It's like Rorschach tests, it's not about what's there, it's about what you see in them. I don't want a bot to make art for me, sometimes I just want some shroom-induced inspirational smears.
- nemomarx 11mo agoI really miss that deepdream aesthetic with the dogs eyes popping up everywhere.
- deleted 11mo ago[deleted]
- csours 11mo agoLOVE IT! It would be really cool if I could zoom out and have everything scale properly!
- Fabricio20 11mo agoWhy is this different per user? I sent this to a few friends and they all see different things from what i'm seeing, for the same time..?
- whimsicalism 11mo agoKimi K2 is obviously the best, but gpt-5 has the most gorgeous ones when it works
- ryandrake 11mo agoI've been struggling all week trying to get Claude Code to write code to produce visual (not the usual, verifiable, text on a terminal) output in the form of a SDL_GPU rendered scene consisting of the usual things like shaders, pipelines, buffers, textures and samplers, vertex and index data and so on, and boy it just doesn't seem to know what it's doing. Despite providing paragraphs-long, detailed prompts. Despite describing each uniform and each matrix that needs to be sent. Despite giving it extremely detailed guidance about what order things need to be done in. It would have been faster for me to just write the code myself. When it fails a couple of times it will try to put logging in place and then confidently tell me things like "The vertex data has been sent to the renderer, therefore the output is correct!" When I suggest it take a screenshot of the output each time to verify correctness, it does, and then declares victory over an entirely incorrect screenshot. When I suggest it write unit tests, it does so, but the tests are worthless and only tests that the incorrect code it wrote is always incorrect in the same ways. When it fails even more times, it will get into this what I like to call "intern engineer" mode where it just tries random things that I know are not going to work. And if I let it keep going, it will end up modifying the entire source tree with random "try this" crap. And each iteration, it confidently tells me: "Perfect! I have found the root cause! It is [garbage bullshit]. I have corrected it and the code is now completely working!" These tools are cute, but they really need to go a long way before they are actually useful for anything more than trivial toy projects.
- fancy_pantser 11mo agoHave you given using MCPs to provide documentation and examples a shot? I always have to bring in docs since I don't work in Python and TS+React (which it seems more capable at) and force it to review those in addition to any specification. e.g. Context7
- ryandrake 11mo agoHaven't looked into MCPs yet. Thanks for the suggestion!
- rossant 11mo agoHave you tried OpenAI Codex with GPT5.1? I'm using it for similar GPU rendering stuff and it appears to do an excellent job.
- deleted 11mo ago[deleted]
- paxys 11mo agoSomething I'm not able to wrap my head around is that Kimi K2 is the only model that produces a ticking second hand on every attempt while the rest of them are always moving continuously. What fundamental differences in model training or implementation can result in this disparity? Or was this use case programmed in K2 after the fact?
- aavshr 11mo agojust curious, why not the sonnet models? In my personal experience, Anthropic's Sonnet models are the best when it comes to things like this!
- xyproto 11mo agoTry adding to the prompt that it has a PhD in Computer Science and have many methods for dealing with complexity. This gives better results, at least for me.
- bigfishrunning 11mo agoWhy does that give better results? Is this phenomena measurable? How would "you have a phd in computer science" change its ability to interpret prose? Every interaction with an LLM seems like superstition.
- xyproto 11mo agoBecause ie. a forum thread that contains this often have better answers, and LLMs are trained on data from the Internet. It's just statistics.
- bpt3 11mo agoIt's wild how much the output varies for the same model for each run. I'm not sure if this was the intent or not, but it sure highlights how unreliable LLMs are.
- eastbound 11mo agoSecurity-wise, this is a website that takes the straight output of AI and serves it for execution on their website. I know, developers do the same, but at least they check it in Git to notice their mistakes. Here is an opportunity for AI to call a Google Authentication on you, or anything else.
- bongodongobob 11mo agoWeird. Sonnet 4.5 one shotted it with: Create an interactive artifact of an analog clock face that keeps time properly. https://claude.ai/public/artifacts/75daae76-3621-4c47-a684-d58654464455 https://claude.ai/public/artifacts/75daae76-3621-4c47-a684-d...
- amelius 11mo agoMaybe they can ask Sora to make variations of: https://slate.com/human-interest/2016/07/martin-baas-giant-real-time-clock-at-schiphol-airport-features-a-man-painting-the-minutes-for-12-hours.html https://slate.com/human-interest/2016/07/martin-baas-giant-r...
- deleted 11mo ago[deleted]
- jcmontx 11mo agoGrok is impressive, I should give it a shot
- Waterluvian 11mo agoHow do they do time without JavaScript? Is there an API I’m not aware of?
- bloppe 11mo agoCSS animation. It's not the real time. Just a hypothetical time.
- Waterluvian 11mo agoI’m imagining some must be using JS because I’m seeing (rarely…) times that are perfectly correct.
- bloppe 11mo agoActually you're right. If you view source, you can see `const response = await fetch(`/api/clocks?time=${encodeURIComponent(localTime)}`);`. I'm not sure how that API works, but it's definitely reading the current time using JS, then somehow embedding it in the HTML / CSS of each LLM.
- vultour 11mo agoIt's crafted with a prompt that gives the AI the current time, then it simply refreshes every minute so the seconds start at zero correctly.
- bhandziuk 11mo agoLooks like css keyframes
- deleted 11mo ago[deleted]
- ssl-3 11mo agoThis really needs to be an xscreensaver hack.
- nasir 11mo agowhere's opus/sonnet! very curious on that!
- ticulatedspline 11mo agoThis is cool, interesting to see how consistent some models are (both in success and failure) I tried gpt-oss-20b (my go-to local) and it looks ok though not very accurate. It decided to omit numbers. It also took 4500 tokens while thinking. I'd be interested in seeing it with some more token leeway as well as comparing two or more similar prompts. like using "current time" instead of "${time}" and being more prescriptive about including numbers
- porphyra 11mo agoLLMs can't "look" at the rendered HTML output to see if what they generated makes sense or not. But there ought to be a way to do that right? To let the model iterate until what it generates looks right. Currently, at work, I'm using Cursor for something that has an OpenGL visualization program. It's incredibly frustrating trying to describe bugs to the AI because it is completely blind. Like I just wanna tell it "there's no line connecting these two points but there ought to be one!" or "your polygon is obviously malformed as it is missing a bunch of points and intersects itself" but it's impossible. I end up having to make the AI add debug prints to, say, print out the position of each vertex, in order to convince it that it has a bug. Very high friction and annoying!!!
- TheKidCoder 11mo agoKinda - Hand waiving over the question of if an LLM can really "look" but you can connect Cursor to a Puppeteer MCP server which will allow it to iterate with "eyes" by using Puppeteer to screenshot it's own output. Still has issues, but it does solve really silly mistakes often simply by having this MCP available.
- firtoz 11mo agoCursor has this with their "browser" function for web dev, quite useful You can also give it a mcp setup that it can send a screenshot to the conversation, though unsure if anyone made an easy enough "take screenshot of a specific window id" kind of mcp, so may need to be built first I guess you could also ask it to build that mcp for you...
- fragmede 11mo agoClaude totally can, same with ChatGPT. Upload a picture to either one of them via the app and tell it there's no line where there should be. There’s some plumbing involved to get it to work in Claude code or codex, but yes, computers can "see". If you have lm-server, there's tons of non-text models you can point your code at.
- pil0u 11mo agoI had some success providing screenshots to Cursor directly. It worked well for web UIs as well as generated graphs in Python. It makes them a bit less blind, though I feel more iterations are required.
- kwanbix 11mo agoWhat a waste of energy.
- mandolingual 11mo agoAlways interesting/uncanny when AI is tested with human cognitive tests https://www.psychdb.com/cognitive-testing/clock-drawing-test https://www.psychdb.com/cognitive-testing/clock-drawing-test.
- hansmayer 11mo agoVery funny. It seems the Qwen generates the funniest outputs :)
- csours 11mo agoOh, Qwen, buddy, you sure are TRYING
- deleted 11mo ago[deleted]
- Imanari 11mo agoQwens clocks are hilarious
- cornonthecobra 11mo agoI like Deepseek v3.1's idea of radially-aligning each hour number's y-axis ("1" is rotated 30° from vertical, "2" at 60°, etc.). It would be even better if the numbers were rotated anticlockwise. I'm not sure what Qwen 2.5 is doing, but I've seen similar in contemporary art galleries.
- gloosx 11mo agoanyone tried opening this from mobile? not a single clock renders correctly, almost looks like a joke on LLMs
- rtcode_io 11mo agoSee https://clock.rt.ht/::code https://clock.rt.ht/::code AI-optimized <analog-clock>! People expect perfection on first attempt. This took a brief joint session: HI: define the custom element API design (attribute/property behavior) and the CSS parts AI: draw the rest of the f… owl
- speedgoose 11mo agoThis is a white page, am I missing something?
- rebelmoon 11mo ago[dead]
- DeathArrow 11mo agoHow can Deepseek and Kimi get it right while Haiku, Gemini and GPT are making a mess?
- 0xCE0 11mo agoSeems like Will's clock drawing test in Hannibal :)
- gwbas1c 11mo agoReminds me of the Alzheimer's "draw a clock" test. Makes me think that LLMs are like people with dementia! Perhaps it's the best way to relate to an LLM?
- hollow-moe 11mo agoobviously they're all broken on firefox, no one uses firefox anyways
- kylecazar 11mo agoNon-determinism at it's finest. The clock is perfect, the refresh happens, the clock looks like a Dali painting.
- jeremycarter 11mo agoLast year I wrote a simple system using Semantic Kernel, backed by functions inside Microsoft Orleans, which for the most part was a business logic DSL processor by LLM. Your business logic was just text, and you gave it the operation as text. Nothing could be relied upon to be deterministic, it was so funny to see it try to do operations. Recently I re-ran it with newer models and was drastically better, especially with temperature tweaks.
- __fst__ 11mo agoThis is why we need TeraWatt DCs, to generate code for world clocks every minute.
- teaearlgraycold 11mo agoQwen 2.5 doing a surprisingly good job (as of right now).
- maxdo 11mo agoSelection of western models is weird no gpt-5.1 , opus 4.1 ( nailed it perfectly ) Something I quickly tested
- accrual 11mo agoI love that GPT-5 is putting the clock hands way outside the frame and just generally is a mess. Maybe we'll look back on these mistakes just like watching kids grow up and fumble basic tasks. Humorous in its own unique way.
- palmotea 11mo ago> Maybe we'll look back on these hilarious mistakes just like watching kids grow up and fumble basic tasks. Or regret: "why didn't we stop it when we could?"
- Bengalilol 11mo agoQwen doesn't care about clocks, it goes the Dali way, without melting. It even made a Nietzsche clock (I saw one <body> </body> which was surprisingly empty). It definitely wins the creative award.
- HarHarVeryFunny 11mo agoLooks like we've got a new Turing test here: "draw me a clock"
- bitwize 11mo agoI'm reminded of the "draw a clock" test neurologists use to screen for dementia and brain damage.
- anon_cow1111 11mo agoI'm having a hard time believing this site is honest, especially with how ridiculous the scaling and rotation of numbers is for most of them. I dumped his prompt into chatgpt to try it myself and it did create a very neat clock face with the numbers at the correct position+animated second hand, it just got the exact time wrong, being a few hours off. Edit: the time may actually have been perfect now that I account for my isp's geo-located time zone
- perfmode 11mo agoi read that the OP limited the output to 2000 tokens.
- lanewinfield 11mo ago^ this! there's a lot of clocks to generate so I've challenged it to stick to a small(er) amount of code
- fouc 11mo agoI wonder if you would get better results if you tell the LLM there's a token limit in the prompt. something like "You only have 1000 tokens. Generate an analog clock showing ${time}, with a CSS animated second hand. Make it responsive and use a white background. Return ONLY the HTML/CSS code with no markdown formatting"
- anon_cow1111 11mo agoI got a ~1600 character reply from gpt, including spaces and it worked first shot dumping into an html doc. I think that probably fits ok in the limit? (If I missed something obvious feel free to tell me I'm an idiot)
- Springtime 11mo agoOn the second minute I had the AI World Clocks site open the GPT-5 generated version displayed a perfect clock. Its clock before and every clock from it since has had very apparent issues though. If you could get a perfect clock several times for the identical prompt in fresh contexts with the same model then it'd be a better comparison. Potentially the ChatGPT site you're using though is doing some adjustments that the API fed version isn't.
- ada1981 11mo agoSonnet 4.5 did this easily https://claude.ai/public/artifacts/c1bb5d57-573b-49e0-9539-71ddda1a8211 https://claude.ai/public/artifacts/c1bb5d57-573b-49e0-9539-7...
- edfletcher_t137 11mo agoLack of Claude is a glaring oversight given how popular it is as an agentic coding model...
- chaosprint 11mo agoThis is such a great idea! Surprisingly, the Kimi K2 is the only one without any obvious problems. And it is even not the complete K2 thinking version? This made me reread this article from a few days ago: https://entropytown.com/articles/2025-11-07-kimi-k2-thinking/ https://entropytown.com/articles/2025-11-07-kimi-k2-thinking...
- esotericwarfare 11mo agoThis is an AD for Kimi K2
- miohtama 11mo agoThe new Turing time test
- bigbluedots 11mo agoI just realized I'm running late, it's almost -2! More seriously, I'd love to see how the models perform the same task with a larger token allowance.
- bigbluedots 11mo agoIs there a "draw a pelican riding a bicycle" version?
- padolsey 11mo agoWe've done this! https://weval.org/analysis/visual__pelican/f141a8500de7f37f/2025-11-09T07-47-57-560Z/compare?prompt=svg-pelican-riding-a-bicycle https://weval.org/analysis/visual__pelican/f141a8500de7f37f/...
- anonzzzies 11mo agoSonnet 4.5 does it flawless. Tried 8 times.
- fouc 11mo agoThe catch was that it was limited to 2000 tokens, i.e. the results get cut off once it hits that.
- fnord77 11mo agowhatever model Cursor uses was telling me the date was March 12, 2023
- imchillyb 11mo agoI love qwen, it tries so hard with its little paddle and never gets anywhere.
- cyberjill 11mo ago666
- superlukas99 11mo ago[dead]
- wanderingmind 11mo agoThe more I look at it, the more I realise the reason for cognitive overload I feel when using LLMs for coding. Same prompt to same model for a pretty straight forward task produces such wildly different outputs. Now, imagine how wildly different the code outputs when trying to generate two different logical functions. The casings are different, commenting is different, no semantic continuity. Now maybe if I give detailed prompts and ask it to follow, it might follow, but from my experience prompt adherence is not so great as well. I am at the stage where I just use LLMs as auto correct, rather than using it for any generation.
- bwhiting2356 11mo agoYou should render it, show an image to the model and allow it to iterate. No person has to one-shot code without seeing what it looks like.
- wewtyflakes 11mo agoIt is funny to see the performance improve across many of the models, somewhat miraculously, throughout the day today.
- stym06 11mo agoIf a human had done this, these would be at a museum
- woopwoop 11mo agoThe qwen clocks are art.
- josfredo 11mo agoWatching these gives me a strong feeling of unease. Art-wise, it is a very beautiful project.
- 3oil3 11mo agoI wonder which model will silently be updated and suddenly start drawing clocks with Audemars-Piguet-level kind of complications.
- jsmo 11mo agolol
- shahzaibmushtaq 11mo agoInteresting idea! Why is a new clock being rendered every minute? Or AI models are evolving and improving every minute.
- Vera_Wilde 11mo agoIt's really beautiful! Super clean UI. The thing I always want from timezone tools is: “Let me simulate a date after one side has shifted but the other hasn’t.” Humans do badly with DST offset transitions; computers do great with them.
- JamesAdir 11mo agoI believe that in a day or two, the companies will address this and it would be solved by them for that use case
- surfingdino 11mo agoWhat a wonderfully visual example of the crap LLMs turn everything into. I am eagerly awaiting the collapse of the LLM bubble. JetBrains added this crap to their otherwise fine series of IDEs and now I have to keep removing randomly inserted import statements and keep fixing hallucinated names of functions suggested instead of the names of functions that I have already defined in the same file. Lack of determinism where we expect it (most of the things we do, tbh) is creating more problems than it is solving.
- anotheryou 11mo agoClaude Sonnet 4.5 with a little thinking: https://imgur.com/a/zcJOnKy https://imgur.com/a/zcJOnKy no thinking: better clock but not current time (the prompt is confusing here though): https://imgur.com/a/kRK3Q18 https://imgur.com/a/kRK3Q18
- themgt 11mo agoJust saw Gemini 2.5 with a little thinking: https://imgur.com/a/nypRD7x https://imgur.com/a/nypRD7x
- arendtio 11mo agoPretty cool already! I use 'Sonnet 4.5 thinking' and 'Composer 1' (Cursor) the most, so it would be interesting to see how such SOTA models perform in this task.
- boxedemp 11mo agoThat's super neat. I'll keep checking back to this site as new models are released. It's an interesting benchmark.
- baidoct 11mo agoGPT-5 looks broken
- Zeraous 11mo agoHow Kımı is better than other BILLION$ companys is really fun
- warpspin 11mo agoLol. This is supposed to replace me at my job already? Great experiment!
- adriatp 11mo agodeepseek representing
- RugnirViking 11mo agowhats going on with kimi k2 and being reasonable/so unique in so many of these benchmarks ive seen recently? I will have to try it out further for stuff. is it any good at programming?
- Bolwin 11mo agoYes, it trades blows with glm for the best open source model
- adi_kurian 11mo agoThink this is just prompt eng tbh. One shot Haiku 3.5 (https://claude.ai/share/66c17968-485e-4d15-974b-4f6958e1e2fd https://claude.ai/share/66c17968-485e-4d15-974b-4f6958e1e2fd) decent looking too. Got it to work on gpt 3.5T w modified prompt (albeit not as good - https://pastebin.com/gjEVSEcJ https://pastebin.com/gjEVSEcJ) `single html file, working analog clock showing current time, numbers positioned (aligned) correctly via trig calc (dynamic), all three hands, second hand ticks, 400px, clean AF aesthetic R/Greenberg Associates circa 2017. empathy, hci, define > design > implement.`
- fouc 11mo agoThe catch was that it was limited to 2000 tokens, i.e. the results get cut off once it hits that.
- adi_kurian 11mo agoAs in, no more than 2K output tokens? Responses above are about ~1K.
- lovegrenoble 11mo agoAre they part of the LLM training set?
- silexia 11mo agoGrok is hilarious
- buzzm 11mo agoWonderful. I don’t particularly care if it is or is not a valid test. I like the “wrong” renderings better. Some are hilarious, some … inspired.