9 ms·
Claude, change the “Add to Cart” button to blue
- johnhamlin 23d agoThis is so accurate it hurts. I’ll be scouring the comments for guys who claim they can’t relate posting links to their magic CLAUDE.mds
- neilellis 23d agoCongratulations!!! You win what’s left of the internet - just ask Claude for your prize! Motrin I’ve had this week. I use codex now.
- satvikpendem 23d agoIt's funny but unrealistic as Claude does a pretty good job at only changing what is required these days with the 5 tier models like Opus 5 or Fable.
- ceejayoz 23d agoThe site is opusfived.com, and Opus 5 is probably the worst so far at doing this.
- rcxdude 23d agoI dunno, in my experience it's often overly-narrow, sometimes jumping through all kinds of hoops to preserve some edge-case behaviour that doesn't matter because I didn't mention it could be changed.
- ceejayoz 23d agoExperiences vary, yes. But you can see in this thread that folks definitely have experienced this.
- brazukadev 23d agoThis is exactly the experience I have with Opus 5. Opus 4.6 is better, Flable 5.1 much better. But Opus 5 is infuriating.
- AIorNot 23d agoLol this is great
- apetresc 23d agoSo this site is just a fan-fiction that thinks it's somehow dunking on Claude? I've never had a session that remotely resembles any of this. I honestly can't tell what point this site thinks it's making.
- edf13 23d ago[flagged]
- ceejayoz 23d agoI wrote my own harness to stop shit like this from getting to my attention out of frustration. I'm sure there's quite a bit of variation from person to person in these sorts of experiences, based on your harness, the way you talk, the stored memory, your CLAUDE.md, etc. But people absolutely have had this Opus 5 style experience the app simulates.
- fg137 23d agoSo you are lucky, congratulations.
- apetresc 23d agoMaybe, if you tried, you could concoct some adversarial example of a stylesheet that a recent Claude model would trip up on and fail to color a button correctly on the first or even second try. But it would have to be some explicitly engineered trick, akin to an optical illusion for humans. You can’t convince me that I’m somehow the odd one out because I regularly have no trouble getting Claude to recolor buttons.
- fractorial 23d agoBrilliant. Precisely the reason I stopped using Anthropic's products.
- deleted 23d ago[deleted]
- appleappleapple 23d agoThis spiked my blood pressure. Well done
- oujiii 23d agoHaha this is spot on how I've been feeling lately. I find it unbearable to work with this model for this reason... any trick out there you can do to steer it not to overcomplicate things? I guess Codex here I come
- ihsw 23d ago[dead]
- RGS1811 23d ago`/model claude-opus-4-7`
- jerf 23d agoThere's also the fact that you are in control. You are not obligated to take the AI's commits. I don't even let it commit much of the time because commit time is review time for me. If it changes the button blue and does four other things, you can just take the blue change and discard the rest. It can't stop you. This isn't a defense of it doing those four other things. It would be nice if it did what you wanted correctly. I'm just saying, as long as our programming skills have not completely atrophied, we have the power. “Ford carried on counting quietly. This is about the most aggressive thing you can do to a computer, the equivalent of going up to a human being and saying "Blood...blood...blood...blood...” ― Douglas Adams, The Hitchhiker's Guide to the Galaxy
- 1970-01-01 23d ago[dead]
- felixgallo 23d agoHere come all the totally organic "wow, I guess I better switch to OpenAI" comments.
- Toutouxc 23d agoThis is so perfect and depressing that I might cry. It’s like a Kafka novel about programming.
- dannypostma 23d agoThis is scary close to my interaction with Claude this week.
- pablopudding 23d agoI’m laughing and crying at the same time. This is what work feels like now. Thank you, well done!
- matsemann 23d agoYeah, I don't mind using AI to help me at work, but having to "talk" with this stupid crap all day will send me to an early pension or something. Can't be healthy in the long run.
- deleted 23d ago[deleted]
- brap 23d agoHow do you manage your frustration in these interactions? I often find myself getting pissed off
- ceejayoz 23d agoThe goal of my personal harness is to get to the point where I never actually talk to Claude directly for that very reason.
- selestify 23d agoIs your harness available for install somewhere? Surely you still have to give feedback to Claude. How do you do that without talking to Claude directly? By using a different model? But wouldn't that AI have no more common sense than Claude?
- ceejayoz 23d agoIt's extremely bespoke. Initial dev required talking to Claude. Now I add a ticket in the board, it makes me a mockup/writeup, I approve, and it gets me a temporary webserver, iOS/Android build, etc. to verify it. Review loops, agents that enforce my pet peeves and testing/debugging processes, etc. all run automatically... and then Codex strips down the prose at the end. There's not zero AI generated output, but it's already been critiqued and verified by a whole cluster of independent actors before it gets to me. When I have feedback, I file a ticket. I wanted to get out of the "what the fuck, why?!" loop. Now I let the agents handle that.
- mrheosuper 23d agoSo instead of talking directly, you talk indirectly through ticket ?
- ceejayoz 22d agoI’m of the feeling that talking directly to these things all day is a path to madness. They talk to each other, come to consensus, and give me a structured thing to look at that has already been vetted and tested and screenshot evidence and a verification plan. It’s not conversational for me.
- andremendes 23d agoI lost it when it finally did the right thing, but then it added a never-requested gradient to the button. Very good!
- fractorial 23d agoI’m impressed you had the patience to even make it that far!
- dwringer 23d agoI only made it through the first round of prompt selection; both options for the second step were equally pointless and not at all prompts I would ever expect to result in a constructive outcome. In my experience, telling the model it screwed up without specifically addressing, unambiguously, how to fix it, only leads to more suffering. If this page illustrates nothing else, I think it shows the immense downside of trying to use simple one or two sentence prompts. EDIT: Actually, I used to use Google's AI Studio a lot and fork it after every successful prompt interaction. When I'd encounter a problematic issue like this, I'd revert to the previous fork and try a different prompt until I could get the desired outcome, thus mitigating the need to "argue" with the LLM. Unfortunately the ability to cleanly fork and revert everything including the LLM context was removed some months ago, and I've yet to discover a workflow with any tool that works as well for me.
- teiferer 23d ago> In my experience, telling the model it screwed up without specifically addressing, unambiguously, how to fix it, only leads to more suffering. I wonder if this is just a reflection of some senior folks being arrogant towards junior folks. When the latter finished a task but not to the liking of the senior person they might just get a "that's wrong, try again". Just to have sth similar repeat the second time around. But the arrogant guy got to boss around the junior one, and some junior folks grow up learning that's how you should behave so they also do it later. Now it's not a person but a machine. And people just make fun of the dumb machine. Well, garbage in, garbage out If you are not specific in what you want, you might get crap back. Or at least sth you didn't envision.
- dazhbog 23d agoPTSD 9000.. I miss the old days, less load bearing BS and more in the zone coding..
- mlekoszek 23d agoFair play. There's a quiet truth to what you're saying, and it's worth pointing out
- burnoutdv 23d agoJust when I came back to my pc and was thinking "I hate this world were everyone talks about AI like fanatics" this made me a little bit happy, especially the unhingend all caps options towards the end
- 8cvor6j844qw_d6 23d agoI find it funny how it went off with subagents and adversarial review when a simple grep or diff is sufficient.
- kaoD 23d agoAm I the only one whose experience doesn't match this? My gripe with Claude is that while investigating how to do this it will report 200 other incidental findings which I overlooked and I realize those are broken too and need urgent fixing, derailing me, not it.
- drcongo 23d agoYou're not alone. I've been sat wondering what kind of codebase someone has if they have this problem, I've never seen this behaviour.
- cub-creature 23d agoOh man, exactly. I'm very prone to scope creep as I work on tasks. I already would notice some things that could be fixed or refactored and have a hard time not touching them before I used agents. But now I have to be very intentional about not letting it manipulate me into fixing EVERYTHING RIGHT NOW. Half the time the "one more thing worth noting, unrelated..." isn't even an actual issue, it just brought it up to fish more usage out of me. Also, while this little demo is certainly exaggerating the issue, I do find working with Claude to sometimes get quite verbose and tiresome. I doubt I would struggle this much to get it to change a button color, but the patterns of speech, the endless lists, the over-explanations, and the whole song and dance of trying to get it to make the change you want without side-effects is frustratingly familiar to me.
- empath75 23d agoClaude is an unbelievable yak shaver if you let it be.
- techscruggs 23d agoI never really understood what being "triggered" was like until now.
- azalemeth 23d agoI've experienced this so many times over. "I was wrong" and "the honest truth" are just forever phrases that are now dead to me.
- ImHereToVote 23d agoI want the dishonest truth.
- column 23d agoUsername may check out
- yomismoaqui 23d agoI don't get the joke... maybe because I'm using Codex?
- dominotw 23d agoits making both buttons blue
- swiftcoder 23d agoThis is pure genius. No notes
- kstenerud 23d agoThat's so weird... This doesn't at all match my experience with Claude. I've never seen it behave this way.
- whalabi 23d agoI've definitely had it behave exactly like this at times and it's infuriating. I think it depends on your codebase.
- inknight 23d agotry using Claude Design
- dominotw 23d agobecause you never changed just one button to blue
- agluszak 23d agoIt means that either you stopped using Claude around Opus 4.6 or you use Fable instead of Opus 5 :)
- kstenerud 23d agoI use Opus 5 for everything.
- teaearlgraycold 23d agoOpus 5 writes too many comments. Other than that I don't agree with what I'm seeing online. It works great.
- raincole 23d agoIt's a joke dude.
- kristofferR 23d agoYeah, but it's not funny since it doesn't match reality.
- jadar 23d agoThis is so good at replicating the experience of frustration, then relief when it finally does what you asked it to do in the first place!
- random_cat_8745 23d agorofl, brilliant
- captainbland 23d agoThis is actually what keeps people using AI: variable reward schedule. It's basically gambling.
- stavros 23d agoPeople say this, but I've never seen it. AI has been very consistent in its rewards for me.
- wuisce 23d agoYou're absolutely right. And it matters.
- mysterydip 23d agoWhich also explains why response speed is so important.
- gbraad 23d agoThis is why I also suspect them to waste tokens on purpose.
- BikiniPrince 23d agoListen Pal, this is load bearing. If you know what is good for you then you will stop asking questions. —Claude
- gbraad 20d ago"Public opinion has been affected. we need a reset. Let me announce it on X."
- Grimblewald 23d agoI really do think its tge case. Ever notice how juuust when you run out of usage its allllmost correct? then the handy helpful "buy more credits!" link pops up. Meanwhile local models are finishing the task without fucking around and without alterting shit it wasnt supposed to touch. How is it a environmental catastrophe level, datacentre requiring, bullshit model is so much worse than something that runs on my workstation and doesn't cost us a ha itable planet? Either they're fucking with us serving 8b models at scale or china really has the AI race in the bag so much so that they can openly release what the USA jealously guards. They moan and complain about china copying from them (while doing the same), but if that's the case in full, then why are the chinese models better? you dont copy bad work and come out ahead.
- deleted 23d ago[deleted]
- deleted 23d ago[deleted]
- andai 23d agoI was expecting it to spend 30 minutes running headless chrome instances, taking screenshots and analyzing them in python to verify the blueness of the result.
- syntaxing 23d ago> 23 agents total. This hit a bit too close to home. Sol has the same issue, spawns a lot of agents for no good reasons (besides burning tokens).
- chrisgarand 23d agoI heard this was a thing when listening to a Theo podcast, he mentioned to add a "Only use subagents if the user explicitly requests them" line in your agents.md file. I don't know if it works, but I've always had a consistent level of token burn on my plans (I've only heavily used Sol after adding it).
- r_lee 23d agonow THAT is a load-bearing simulation
- dsign 23d agoThat was funny :-) I use Opus and Sonnet 5 all the time and I find their language grating. But honest, I prefer to put up with it and get the results than to put up with my own human limitations and not get the results.
- chrismorgan 23d agoI’ve never used any of these tools. Please tell me that this is a grossly exaggerated parody, and that the tools don’t write like this, or do so many ridiculous things. For my sanity. (I am genuinely uncertain, though I presume it’s at least somewhat exaggerated.)
- dd8601fn 23d agoNo. It’s just a bunch of jokes rolled up into a big exaggeration. It’s funny because there are elements of truth in each bit of it, though.
- retsibsi 23d ago> Please tell me that this is a grossly exaggerated parody, and that the tools don’t write like this, or do so many ridiculous things It's a pisstake, but (in the bits I read, and based on my own personal experience) the writing style is barely exaggerated, while the behaviour doesn't ring true at all.
- selestify 23d agoThe behavior, while slightly exaggerated, rings entirely true for me. From the other comments in this thread, it seems I am prompting poorly in a similar way to the options offered. I am guessing you prompt differently than what is shown in the game?
- deaux 22d agoI'm not the person you were replying to, but yes. Let me show you the issue. "Why is half the site blue now? I asked you to change one button."; half the site isn't blue, why would you say that? What would "half the site" even mean? Ironically even in this satire supposed to make fun of how Claude responds to prompts, it is smart enough to ignore that. A better prompt would be "Why are many parts of the site, such as the Cancel button, now also blue?", though in this case it doesn't matter as the reply would be roughly the same. Next, Claude helpfully tells us we're dealing with a monstrosity of a codebase out of hell: "Seven components consume the token directly; another eleven reach it through aliases; three use it only in hover/focus state; and two appear to be accidental cross-role consumers" is such an insane codebase that the initial request was basically impossible to carry out. Claude made that perfectly clear, yet the next prompt choice just ignores that fact. "Revert everything except Add to Cart. That is the whole task." It just explained why this is impossible. It had only made one change, so there is no "revert everything except X"; there's only a single change to revert, which would mean Add to Cart too would no longer be blue. The human just chose to ignore that. The other option is "I do not care about the token architecture. I do not care why it happened. Put everything back the way it was and leave ONLY the Add to cart button blue. Please.". This has the same issue; "Put everything back and _leave_ only the button blue" is an impossible ask, which it just explained to you. So if you really don't want to deal with cleaning up some of the mess, instead only adding on to it but getting your blue button, then you'd want to say "Revert the change you made, then make a new change that scopes the new blue color to only the Cancel button". Ignoring the information it gives you, or telling it things that are incorrect, of course will lead to bad results. It's impossible not to.
- vant 23d agoglad to see I'm not the only one... anthropic needs to support my anger management treatment
- airstrike 23d agoThis does not match reality at all, speaking as the #1 user on agent hours per clauderank.com
- alentred 23d ago-= CAUTION, SPOILERS =- This got me on "cyanide blue", and I was ROLLING ON THE FLOOR LAUGHING on "Approaching usage limit". I can barely stop laughing now and my stomach hurts. I mean, Thank You!
- 100percentjake 23d ago"Confirming the button contains no cyanide" sent my sides firmly into orbit. This is fantastic.
- tamimio 23d agoThis is gold, thanks for the giggles! I think it was designed that way to burn tokens.
- GracefullyShot 23d agoit gave me headache in 2 turns, just like opus 5 !
- almostdeadguy 23d agoNails the Claude dialect. Technical nonsense like: > I'm collapsing this back to the rendered outcome: And intermixed with SaaS product page idioms from a brain-damaged marketer like: > No broader cleanup. > No further architecture work. > Just the button. Aside from the patterns everyone knows like em-dashes, "its not X, it's Y", etc. I think the key features of claude diction is it sounds like a junior engineer over their skies who is trying to make up for that with extra verbiage mixed with extremely grating SaaS marketing-ese.
- totetsu 23d agoI was waiting for it to .. say usage limit reached after reverting it back to how you started..
- g-b-r 23d agoIt reaches the usage limit after adding a cookie banner above the button
- telesilla 23d agoThis felt very late-90s net art. Stressful but nicely done satire.
- deleted 23d ago[deleted]
- dwedge 23d agoI got way too annoyed at this before realising it was an optional game and I could just close the tab
- alex_c 23d agoSurprising how much of life this applies to when you really think about it!
- btown 23d agoHacker News is the epitome of this! If you find yourself not enjoying your daily dose of "someone is wrong on the internet" (via the immortal https://xkcd.com/386/ https://xkcd.com/386/) as you find yourself crafting the perfect response, you can always close the tab!
- SoftTalker 23d agoI actually don't end up clicking the "reply" button on a good portion of the replies I start to write.
- barbazoo 23d agoI was gonna respond to you but actually decided not to.
- jodrellblank 23d agoDon't worry, Reddit/Facebook/Gmail et al. still transmit that draft reply to their servers, store it against your profile, and train on it. Probably.
- iririririr 23d agoand ironically, most of them don't even offer the draft feature! Facebook, tiktok, etc... they send your typed message, but if the app crashes or you close it, there's no draft anywhere you can find ;)
- gwbas1c 23d agoI don't get who this is making fun of: - The people who won't make any effort to learn the tools, and something as simple as reverting code (via git) needs to be done by AI? - The awful programmers who we've had to endure working with, who are so bad at simple changes that they have negative productivity? - Or Claude itself? --- BTW: I don't have these problems, but I'm also not afraid to do things myself when it's easier. Edit: If I want to change a button's color, I just change it manually. If I don't know where the code for the button is, I might start with prompting, (because AI can often find the code faster than I can,) and then once the diff is proposed, start adjusting things by hand.
- kmoser 23d agoIt's reductio ad absurdum, satirizing the Claude experience.
- gwbas1c 23d agoThen, IMO, the joke is lost: This feels like working with various forms of difficult, immature, indecisive, engineers; and possibly with very disorganized codebases where small changes require major refactoring.
- mrheosuper 23d ago>very disorganized codebases where small changes require major refactoring. In the vibecode era, this is more common than you thought.
- gwbas1c 22d agoI don't think it's more or less common, pretty much any code base I inherited prior to AI had serious problems.
- kmoser 23d agoAre you saying you've never had an experience even somewhat like this while working with Claude, or any LLM for that matter?
- dmd 23d agoWas this made by someone who hasn't actually used any of these tools in over a year?
- _fat_santa 23d agoAt least with Codex, this has not been my experience at all. It still screws up sure, but in every case I can ask "why did you do this" and it can trace back what made it take that particular decision. Typically it's always that I either didn't specify the problem correctly or made a really dumb mistake (executing the task on the wrong project....did this one yesterday) or it's something within a skill file that instructs it (at which point I fixup the instructions). Once in a blue moon it's actually the model making a material error in it's thinking and I have to go back and redo it.
- arnorhs 23d agoagreed to some extent. I think this parody still highlights what I feel is often the experience. It might not happen on a simple task such as changing a button color, but on more complicated things, this can definitely be exactly what it feels like.
- tarxzvf 23d agoModels hallucinate plausible answers to why they did things. It might be true and it might be complete fiction.
- taeric 23d agoI'm growing increasingly confident that this is how people often work, as well.
- skinfaxi 23d agoPeople don't make rational decisions that make rationalized decisions. Is there any thought to pulling your hand off a hot surface?
- autoexec 23d agoPeople do both. Some choices aren't worth the time and effort of detailed analysis and contemplation and some are basically instinctual, but there are plenty of times that choices are carefully considered and well reasoned before being made and acted on.
- deleted 23d ago[deleted]
- inerte 23d agoTo be fair I’ve worked on human programmed systems where similar “it should be a half point story” requests would be met with snark by the engineers and take 2 sprints. I guess we are all PMs now.
- inerte 23d agoAlso likely, devs took shortcuts to deliver fast. Now to make the button blue they need to differentiate primary buttons from others. Simple, right? But design guidelines prevent one offs, and no !important. So you create a CSS class, but you discover another element on the header declared itself as primary (the search icon or the sign in button). You talk to that team and they decided to scope what’s primary according to their component. To change the sign in button to grey now you need to talk with the growth team. Growth team wants to run an experiment but they’re backlogged, only next quarter. They say you can innersource, just need VP approval. VP says blue matches a marketing campaign that is about to go out, agency has already been hired. You can’t talk to the agency unless Legal approves. So you leave the button gray, to revisit decision next planning cycle after you can align all stakeholders.
- mring33621 23d agoJust understand that every requested change results in a game of whack-a-mole.
- deleted 23d ago[deleted]
- deleted 23d ago[deleted]
- vinc 23d agoYou should plan the task before implementing it to make sure that it will do the right thing.
- akho 23d agois this using my subscription
- jmartrican 23d agoWow that gave me anxiety... lol. Ok cool so I'm not the only one who gets into these situations.
- Nevermark 23d agoFunny exercise. For a moment I thought, wow, someone put a lot of work into creating this theme park of frustration. Next: It would be so easy to create a faux-Claude like this. Then: How hilarious to watch the transcripts of unsuspecting users in real time. Finally: I began wondering if this might be relevant to all the redundant, unnecessarily preambled, sentence structure complexifying, indirect referencing, canned phrasing, ambiguity mining, analogy maxxing, over-wordy responses I have recently been getting from Fable...
- arbirk 23d agoOne thing I have to be honest about, and it's mine.. The one thing I would check before... do you want to do that? Say go an and will do it without the check While checking I found 3 vulnerabilities and 2 potential optimizations of which I fixed 2 and 1. Do you want me to file the other as issue, or stop for the day? We have done <lists a weeks worth of work> this morning. I feel you need a break
- improbableinf 23d agoThank you for creating this. Just thank you
- deleted 23d ago[deleted]
- moralestapia 23d agoWow, this is SO on point. It made me stop using Claude at all. Codex has almost surgical precision, and I like that a lot. (But nowadays I just use DeepSeek Flash. it does screw up but its cents so ¯\_(ツ)_/¯).
- Kim_Bruning 23d ago1970-01-01's Kobayashi Maru solution is the only thing that gave me closure :-P but unfortunately it's [dead].
- mzajc 23d ago> "`#16b8c4`. Yes. Apply it." > WebFetch en.wikipedia.org/…/Cyan > WebFetch en.wikipedia.org/…/Prussian_blue > WebFetch www.colorhexa.com/16b8c4 Brilliant.
- pohl 23d agoAmusing, but do people actually prompt in the style of any of the options given? All this for what is ultimately a PEBCAK error.
- bbstats 23d agoMine immediately did it correctly?
- matthieu_bl 23d agoConfirmed, Anthropic are silently testing Mythos 5.1 with some users
- jezzamon 23d agoFunny game. Do people really prompt AI like this? Multiple times the choice was either to yell at the agent, or ask it why it did something, neither of which are very fruitful lines to go down if you know what you're doing
- bennettpompi1 23d agothis is hilarious lmao
- johnisgood 23d ago> Why is half the site blue now? I asked you to change one button. > Half the site is blue. I asked for ONE button. Those are my only options when the site is clearly not blue, two buttons are. There is a reason for why I am much more specific than this.
- plorkyeran 23d agoYeah, if this is how people interact with claude I’m not surprised they’re having a bad time in ways that I don’t. Asking it why it did something or getting combative is a waste of time.
- jaggederest 23d agoClear context, revert the commit, change the prompt or documentation, try again.
- selestify 23d agoHow does one learn to interact with Claude more effectively?
- tsukikage 23d agoRealise that you are talking to some mathematics in a box. You can't get a rise out of it. It cannot feel guilt or remorse. Whatever emotional payoff one might want from "I asked for ONE button" cannot be had here. Meanwhile, not only are you being charged by the word, but every word you send that is not directly on the path to getting what you want done is just noise in the maths getting in your way. So leave emotion at the door and make your words count. Voice your frustrations at the stress doll next to your monitor, sure, but spending tokens on them us just wasting time and money. Explicitly typing out your exasperation will not get you closer to whatever it is you are trying to accomplish. What can you type that will? Laser focus: clearly, concisely explain what is wrong right now, and how you want it solved. Then do that again until all the moles have been whacked. If the AI is stuck in a sycophancy loop, start a new clean session. Do this frequently anyway: every turn in the same session charges for all preceding conversation again /and/ degrades the LLM's performance.
- NickNaraghi 23d agoYou didn't say please or thank you.
- snkline 23d agoSeems to be getting a polarized response. I quite enjoyed the it, but I do think the creator should have made it clearer that a) it is in fact a joke site and b) it does not consist of actual Claude responses. It is easy to misinterpret this site, and therefore not "get" the joke.
- deleted 23d ago[deleted]
- deleted 23d ago[deleted]
- khernandezrt 23d agoI mean honestly if you're using an agent for something this simple you deserve this and all the token usage that comes with it.
- bdelmas 23d agoIt would have been funny a year ago but now I have no issues of that sort or even for more complex tasks
- chuckadams 23d agoNot really my experience with Claude, and the prompts are not how I would phrase them, but still pretty damn funny: I especially loved the slot-machine-style picker for LLM-isms.
- dudeinhawaii 23d agoGreat site, triggered memories! haha. To try to add something to this discussion -- I think that while I've seen these sort of loops less --- what I have seen is "overly helpful". Models nowadays want to double-triple-quadruple check things. I'm being silly but it verges on "I have a working solution but let me write a variation in Rust to ensure a convergent solution and prove this works". I've had to stop models nowadays mostly because they're being agonizingly pedantic in their validation. Opus is actually one of the most pedantic and "off track" here. But again, not in a bad way. I'm usually like "stop testing latency between 50 runs of this app... this is version one.. we're going to make a million more changes.. you're not buying us anything".
- mrinterweb 23d ago> triggered memories Yeah of 10 minutes ago. It is shocking how long some seemingly simple things can take. I know there are some things I can do faster than the LLM and some things it can do faster than me. The amount of rambling BS is the exhausting part.
- epistasis 23d agoTry different models, it's a breath of fresh air. GLM 5.2, etc. all make life much more enjoyable. They may not one-shot a complex project the same way that Claude can spit out memorized architectures, but that sort of system is always only useful for a one-off prototype anyway, so not much is lost. Edit: for how to do this, I set up an OpenRouter account so that I could easily switch models, and then ran them in Pi inside of Orca ADE. Orca lets me easily switch from Pi to Claude Code to Kilo to Codex or Hermes or whatever. Pi+OpenRouter lets me easily switch the LLM. All of it lives in a single open source orchestrator to avoid any platform lock in to any AI company going forward, and can even do local LLM should I care to take on that massive hassle.
- torginus 23d ago> Yeah of 10 minutes ago. It is shocking how long some seemingly simple things can take. And shocking how little code end effort some things take if done by hand. I am not some hardline LLM hater, just venting my frustration.
- cropcirclbureau 23d agoIs my job a joke to you??
- arrowsmith 23d ago[dead]
- nullbio 23d agoThanks, I hate it.
- robinpie 23d agoIf you interact with Claude like this and ignore legitimate issues it flags, no wonder you have a bad experience.
- selestify 23d agoWhat legitimate issues are there around making a fucking button blue?
- gitowiec 23d agoThis is kind of funny but with tears (not off joy). I stopped playing because it made me angry
- homeonthemtn 23d agoLord this triggered my eye twitch
- Culonavirus 23d agoThis is the smoking gun. Yes, and it's mine!
- JohnMakin 23d ago> Worth naming: the Add to cart button is still black. Got an audible guffaw out of me. This really is what the experience is like sometimes if you're just giving it a result without being specific in implementation, and it comes out of nowhere, some days much worse than others. I've become patient with it, but whatever this style of output is called or doing - it is both condescending and entirely unhelpful, and it seems designed to frustrate.
- GreenWatermelon 23d agoI've come to call this language "Claudese English" (or Claudish) Recently I've started using GLM models and noticed they aren't very claudish (just a little bit, compared to DeepSeek which is extremely claudish)
- rcfox 23d agoA Bot & Costello
- xd1936 23d agoLaughed out loud at the overly cautious Terms of Service that it generated for "Cyanide Blue", the color it made up
- stevenhubertron 23d agoYou can hate AI all you want, but this is because of a bad design system, not because of a bad LLM.
- deleted 23d ago[deleted]
- monooso 23d agoOh god, it's so painfully accurate.
- qazxcvbnmlp 23d agoOh dear - this explains why people have bad experience with ai. The prompts in the game are terrible. The user was providing no context, they had no ability to give the model background or context. A simple why would have prevented 95% of these side quests. “Im trying to increase the relative visibility of the add to cart action on the page. Can we please change it to blue without changing any other buttons. /effort low. Let me know if you have questions and before editing anything tell me what you are going to do”
- K0IN 23d agonow add a source tab and lets see, how many ppl. will fix it themselves and how many turns it needs.
- epistasis 23d agoOne note for those still using the Claude system for chats: there's no system to get generated images, spreadsheets, etc. out of the system. They claim it's a "security concern" to provide that data to you, as if they are protecting you by refusing to follow data export laws. I'm hesitant to email their data emails, as it's common for companies to delete all data upon any request, instead of providing data as they are required to.
- josh_p 23d agoa funny codex anecdote: I had 5.6-Luna coordinate a code review in which it spawns 2 agents looking for different things. My prompt was "review the currently checked out branch. diff target is `next`. The jira ticket is XX-XXXXX..." My `next` branch was a few commits behind `origin/next` but it still did its review against the stale local version instead of clarifying or inferring that I meant `origin/next`. The findings were very confusing until I realized what I did. I'm noticing the need to be really specific with any instructions lately, which I don't think is a bad thing. I expect co-workers (or anyone really) to tell me what they need in specific terms so I can get it right. I can do the same for the machine, I guess.
- teekert 23d ago[dead]
- heaney-555 23d agoI haven't experienced anything like this with Codex. Why do people stick with Claude Code if it doesn't do what you ask it to?
- deleted 23d ago[deleted]
- jwrallie 23d agoIt’s a feature not a bug kind of scenario. People will also praise when it reads the code and do something out of “common sense” that happens to align with what you want. I enjoy working one level lower where I describe the files I want changed and how it should be solved, I either have bad results or stop understanding the structure of the programs otherwise. I also would never go more than two prompts in without reverting and starting differently nowadays.
- deleted 23d ago[deleted]
- digitaltrees 23d agoI am self delusional enough to think I could have gotten a blue button if I could have typed the prompt. But scared enough to realize I’ve lived this and spent months refactoring these exact problems. Hold on let me go commit code without reviewing.
- GoToRO 23d agoThis is not a game. Choosing from wrong answers only is not a game.
- nonethewiser 23d agoI dont find this to be indicative of Claude (opus?) at all. My experience doesnt lead me to think it would change a cancel button to blue if I ask it to change an "Add to Cart" button to blue. I assume this is just a contrived example?
- paimapi 23d agoI think it's best read as a humorous piece of creative fiction and you can employ your suspension of disbelief for this ride
- nonethewiser 23d ago>I think it's best read as a humorous piece of creative fiction and you can employ your suspension of disbelief for this ride But what is the point of the fiction, if not that its relatable? Is it supposed to mirror some fictional reality that the author is relieved we dont live in? Or is it supposed to parody real life? I think its supposed to parody reality, but I don’t see the resemblance.
- data-ottawa 23d agoIt’s a contrived example, but also not very far from the mark. The spinner with random claudisms is what makes the game, mixed with Claude helpfully deciding to pick up some tasks along the way. I feel like most of my conversations with Claude lately are a battle of “I am a technical user, share technical details, but not random superfluous gibberish”
- nunez 23d agoWow this is EXACTLY my experience building a CLI with Claude, but I think with Opus 4.6. I gave it a README specifying the requirements for everything I wanted to build. My interaction with Claude was more or less like this.
- marktl 23d agoRotflol!
- ricardobeat 23d ago[flagged]
- bethekidyouwant 23d agoEven funnier is all the ways people are malding about this in the thread.
- anthomtb 23d agoI’m don’t do much change-clicky-button software development. But when I do, I point models to specific lines of code. And have never had a result this bad. Maybe this is geared towards pure vibe coders for whom a file and line number is too technical.
- deleted 23d ago[deleted]
- atleastoptimal 23d agoThese things are cute but they basically become outdated in a few months as the models get better.
- satirev 23d agoThis is elite satire
- sergiotapia 23d agoI only lasted two turns before I had to close the tab. I can't stand anthropic models.
- charcircuit 23d agoI bet if you used the real Opus 5 it could one shot this.
- butterNaN 23d agoI wanted to see what the various phrases in the "slot machine" are. Surprised it doesn't have "LOAD BEARING" in there somewhere: Caveats: • ONE HONEST CAVEAT • ONE THING WORTH STATING PRECISELY • ONE THING WORTH FLAGGING • WORTH NAMING • WORTH STATING PLAINLY • I DON'T WANT TO LEAVE THIS IMPLICIT • I'D BE DOING YOU A DISSERVICE • I DON'T WANT TO BURY THIS • BETTER NOW THAN LATER • ONE HONEST TRADEOFF • I DON'T WANT TO PAPER OVER THIS • ONE SMALL HOUSEKEEPING ITEM • THE HONEST PART IS SIMPLER Pushbacks: • FAIR PUSHBACK • FAIR HIT • THAT'S ON ME • YOU'RE RIGHT ABOUT THAT • YOU'RE RIGHT • YOUR INSTINCT IS RIGHT • FAIR, AND MORE SPECIFIC THAN IT SOUNDS • RIGHT, FOR A REASON WORTH NAMING • I'M NOT GOING TO DEFEND THAT • YOU'RE RIGHT TO PUSH BACK Reframes: • LET ME BE PRECISE • THE SHARPER DISTINCTION • THE PART THAT MATTERS • THE USEFUL PART IS NARROWER • LESS X THAN Y • VISIBLE ISSUE / UNDERLYING ISSUE • THOSE ARE DIFFERENT CLAIMS
- deleted 23d ago[deleted]
- deleted 23d ago[deleted]
- commandlinefan 23d agoIs this a problem with Claude updating code that a human wrote, though? Would Claude do better on code that it started on its own? Humans have a bad tendency to write unmaintainable code (usually at the behest of managers breathing down their necks to HURRY UP even when it doesn't matter). If the code had been designed with good coding practices from the beginning, I wonder if Claude would have struggled so much with it.
- herrherrmann 23d agoThat implies that Claude writes better code by default, which isn’t necessarily true. Especially as projects get bigger, you can easily end up with unmaintainable LLM-written code if you don’t actively intervene and know a cleaner way.
- annoyingnoob 23d agoI think I have PTSD after that.
- nelaggy 23d agodelightful user experience 10/10
- stillpointlab 23d agoI get the joke, but none of the offered prompts are close to how I speak with coding agents. I felt like I was being forced to feed garbage into the machine and then I'm supposed to act surprised when garbage came out.
- wh0ami 23d agowhat a interesting news!
- stefanhall05 23d agohahah awesome I love it
- almosthere 23d ago[dead]
- jebarker 23d agoThat triggered a physical response of tightening in my stomach and heavier breathing due to frustration.
- moritzwarhier 23d agoI love this. Is there a real generalized and capable, but frustratingly "evil" coding model? It could fix so many more buttons in no time, and ideally open follow-up issues about the open questions.
- classified 23d agoThis is very instructive. Now I finally understand why so many developers claim they are much more productive with AI coding. Merely writing code without adversarial fights is just not load-bearing enough. How could we ever live without this before?
- wvbdmp 23d agoWhat shared style? Each button has its own extremely verbose style attribute specifying, among many other things, custom transition curves.
- aand16 23d agoThis is spiking my cortisol
- smb06 23d agoThis is genius. Did you use Claude to build this?
- danwitt 23d agoPersonally I want to see the obnoxious comments it’s leaving to poison future sessions.
- mindcrash 23d agoI just had a really interesting yet (funny) conversation with ChatGPT (5). We came up with this idea to make a custom "light box" for two Govee COB light strips which would enable both wall washing (leds shining towards the wall, to the top) and lighting up the wall panels and the desk below the box (leds shining to the bottom) using two led profiles tucked away in this construction. Text wise it seemed it more or less understood exactly what I wanted... ... but then I asked to draw a simple ASCII layout so my DIY savvy but not really that strong in English dad could understand better what the plan was. But every single time it got the orientation of the led profiles wrong. I told it the orientation was wrong. It understood that the orientation was wrong (it kept saying sorry even) but even when warned multiple times the orientation was still wrong. I really wonder if Astra will be smarter than this, and if not AGI still has a veeeeery long way to go...
- bartread 23d agoThis is an amazing piece of satire but it's really not so far from reality. I've lost count of the number of times I've submitted a prompt to Claude, had it interview me about areas of uncertainty, iterated on the plan a handful of times and then had it one shot the generation of 1000 - 3000 lines of code and tests that work pretty much perfectly first time but then had it chew through six figures of tokens and achieve absolutely nothing useful at all on the seemingly simplest of tasks. This site is clearly based on bitter and hard won experience with the real product.
- dbg31415 23d agoToo soon.
- emilfihlman 23d agoIs this what it feels like using Claude? Holy shit this was the most infuriating thing I've experienced in a while lmao, I felt so relieved when the button was turned to blue, but then so visceral "no no no no no" panic when it turned to the gradient. Amazing E: this is a horror simulator, my breathing is becoming shallow and fast, I love it.
- fortran77 23d agoI have found that "trivial" UI stuff is the hardest for AI to get right.
- hk1337 23d agoIt's like the scene from Liar Liar, him trying to say the blue pen was red.
- basilgohar 23d agoThis...this is giving me PTSD...
- swordsith 23d agoFirst thought looking at the first options was I would absolutely never prompt this way. The LLM may as well be a misaligned genie in a bottle, it has to be treated as such.
- BobbyTables2 23d agoThis is worse than asking an intern to fix a shipping product…
- mrheosuper 23d agoI lose it after removing the ToS and the button becomes gray lol.
- DimmieMan 23d agoThis is fun and hits too close to home, I can appreciate that Claude won't ever screw up this badly for something this simple but it's a simplification that captures the essence of the rage loop perfectly. It's making me realise something new about why I find Claude so exhausting. Of course you're getting terrible results with these prompts but I found myself getting just as enraged contemplating what my reply would be. Estimating every single way this little asshole will hyper-focus on a pointless bit of semantics, ignore explicit directions, go do something unrelated and spin its wheels until you’re out of credits or wasted a ton of time. The defensive writing is honestly just as exhausting as reading Claude’s output. Every session is a game of Russian roulette with a chance you'll end up in one of these loops. At that stage all the joy you had of having it pump along with your intentionally crafted prompts and agent configuration is undone in an afternoon.
- mrheosuper 23d agoPeople keep saying those prompts are unrealistic, and no one would prompt like that. But those prompts are exactly what my managers/product people would use when they decide to do it themself after firing the entire dev department.
- nirmeet011011 23d ago[flagged]
- ovasoncn 23d agoIf you're a programmer, the blue-button test is incredibly annoying. But then you remember this same model is deployed in Claude Gov. If a government employee says, “Close one exit of the New York subway,” and Claude responds with the same annoying “before I do that, here are the downstream consequences you may not have considered,” you suddenly understand why this behavior exists.
- deleted 23d ago[deleted]
- SamInTheShell 22d agoIf this is anyone’s experience with Claude for a button today, I feel like Jobs would say you’re holding it wrong.
- weihz5138 22d ago[dead]
- weihz5138 22d ago[dead]
- jrobertgardzins 22d agoWall of text to just change the button color! Can't believe it. Gives a good feeling of vibe coding .
- deaux 22d agoThe annoying there here isn't Claude, it's the "human". "It's not X but Y"-satire not intended. It was annoying to be forced to send such dumb replies. Half the site isn't blue, why would I say that? What would "half the site" even mean? Claude helpfully tells us we're dealing with a monstrosity of a codebase out of hell: "Seven components consume the token directly; another eleven reach it through aliases; three use it only in hover/focus state; and two appear to be accidental cross-role consumers" is such an insane codebase that the initial request was basically impossible to carry out. Claude made that perfectly clear, yet the next prompt choice just ignores that fact. "Revert everything except Add to Cart. That is the whole task." this is impossible. It just explained why. You just chose to ignore that.
- vrighter 22d agoThis made me laugh out loud. And got my office coworkers to give me a strange look. Thank you
- moezd 22d ago1) If your LLM is behaving like this, you are imprecise in your input. You can also stop responding back and just edit your previous input with the "wisdom" that it shares with you in the current iteration. 2) Don't argue with AI. If you end up in a place where you're actually losing argument against it, stop. Ask it to steelman your position. Let it unwrinkle itself.
- spncai 22d ago[flagged]
- digitalhobbit 21d agoJust came across this and came here to post it. As much as I love Claude Code, this satire is quite accurate.
- lucas_t_a 19d agoi like the decision prompts taking the whole fucking screen, so true.