10 ms·
Recently, I asked Codex CLI to refactor some HTML files. It didn't literally copy and pasted snippets here and there as I would have done myself, it rewrote the
by rossant 1y ago
Recently, I asked Codex CLI to refactor some HTML files. It didn't literally copy and pasted snippets here and there as I would have done myself, it rewrote them from memory, removing comments in the process. There was a section with 40 successive <a href...> links with complex URLs.
A few days later, just before deployment to production, I wanted to double check all 40 links. First one worked. Second one worked. Third one worked. Fourth one worked. So far so good. Then I tried the last four. Perfect.
Just to be sure, I proceeded with the fifth one. 404. Huh. Weird. The domain was correct though and the URL seemed reasonable.
I tried the other 31 links. ALL of them 404ed. I was totally confused. The domain was always correct. It seemed highly suspicious that all websites would have had moved internal URLs at the same time. I didn't even remember that this part of the code had gone through an LLM.
Fortunately, I could retrieve the old URLs on old git commits. I checked the URLs carefully. The LLM had HALLUCINATED most of the path part of the URLs! Replacing things like domain.com/this-article-is-about-foobar-123456/ by domain.com/foobar-is-so-great-162543/...
These kinds of very subtle and silently introduced mistakes are quite dangerous. Be careful out there!
- ivape 1y agoYou’re just not using LLMs enough. You can never trust the LLM to generate a url, and this was known over two years ago. It takes one token hallucination to fuck up a url. It’s very good at a fuzzy great answer, not a precise one. You have to really use this thing all the time and pick up on stuff like that.
- doikor 1y agoI would generalise it to you can’t trust LLMs to generate any kind of unique identifier. Sooner or later it will hallucinate a fake one.
- wat10000 1y agoI would generalize it further: you can't trust LLMs. They're useful, but you must verify anything you get from them.
- grey-area 1y agoOr just not bother. It sounds pretty useless if it flunks on basic tasks like this. Perhaps you’ve been sold a lie?
- ivape 1y agoWell, you see it hallucinates on long precise strings, but if we ignore that, and focus on what it’s powerful at, we can do something powerful. In this case, by the time it gets to outputting the url, it already determined the correct intent or next action (print out a url). You use this intent to do a tool call to generate a url. Small aside, it’s ability to figure what and why is pure magic, for those still peddling the glorified autocomplete narrative. You have to be able to see what this thing can actually do, as opposed to what it can’t.
- sebtron 1y ago> Well, you see it hallucinates on long precise strings But all code is "long precise strings".
- ogogmad 1y agoHe obviously means random unstructured strings, which code is usually not.
- grey-area 1y agoI can’t even tell if you’re being sarcastic about a terrible tool or are hyping up LLMs as intelligent assistants and telling me we’re all holding it wrong.
- IanCal 1y agoThey're moderately unreliable text copying machines if you need exact copying of long arbitrary strings. If that's what you want, don't use LLMs. I don't think they were ever really sold as that, and we have better tools for that. On the other hand, I've had them easily build useful code, answer questions and debug issues complex enough to escape good engineers for at least several hours. Depends what you want. They're also bad (for computers) at complex arithmetic off the bat, but then again we have calculators.
- hansmayer 1y agoYeah so, the reason people use various tools and machines in the first place is to simplify the work or everydays tasks by : 1) Making the tasks execute faster 2) Getting more reliable outputs then doing this by yourself 3) Making it repeatable . The LLMs obviously dont check any of these boxes so why don´t we stop pretending that we as users are stupid and don´t know how to use them and start taking them for what they are - cute little mirages, perhaps applicable as toys of some sort, but not something we should use for serious engineering work really?
- ivape 1y ago[flagged]
- exe34 1y agoI take it you have bought stocks? What do you recommend?
- jsjshxnsbs 1y agoI'll take the L when llms can actually do my job to the level I expect. Llms can do some of my work but they are tiring they make mistakes and they absolutely get confused by a sufficiently complex and large codebase. Quite frankly, not being able to discuss the pros and the cons of a technology with other engineers absolutely hinders innovation. A lot of discoveries come out of mistakes. Stop being so small minded.
- laterium 1y agoWhy is the bar for it to do your job or completely replace you? It's a tool. If it makes you 5% better at your job, then great. There's a recent study showing it has 15-20% productivity benefits: not completely useless, not 10x. I hope we can have nuance in the conversation.
- hansmayer 1y ago...and then there was also a recent MIT study showing it was making everyone less productive. The bar is there because this is how all the AI grifters have been selling this technology - no less than end of work itself. Why should we not hold them accountable for over-promising and under-delivering? Or is that reserved just for the serfs?
- jollyllama 1y ago> You’re just not using LLMs enough. > You can never trust the LLM to generate a url This is very poorly worded. Using LLMs more wouldn't solve the problem. What you're really saying is that the GP is uninformed about LLMs. This may seem like pedantry on my part but I'm sick of hearing "you're doing it wrong" when the real answer is "this tool can't do that." The former is categorically different than the latter.
- IanCal 1y agoIt's pretty clearly worded to me, they don't use LLMs enough to know how to use them successfully. If you use them regularly you wouldn't see a set of urls without thinking "Unless these are extremely obvious links to major sites, I will assume each is definitely wrong". > I'm sick of hearing "you're doing it wrong" That's not what they said. They didn't say to use LLMs more for this problem. The only people that should take the wrong meaning from this are ones who didn't read past the first sentence. > when the real answer is "this tool can't do that." That is what they said.
- jollyllama 1y ago> If you use them regularly you wouldn't see a set of urls without thinking... Sure, but conceivably, you could also be informed of this second hand, through any publication about LLMs, so it is very odd to say "you don't use them enough" rather than "you're ignorant" or "you're uninformed". It is very similar to these very bizarre AI-maximalist positions that so many of us are tired of seeing.
- IanCal 1y agoThis isn't ai maximalist though, it's explicitly pointing out something that regularly does not work! > Sure, but conceivably, you could also be informed of this second hand, through any publication about LLMs, so it is very odd to say "you don't use them enough" rather than "you're ignorant" or "you're uninformed". But this is to someone who is actively using them, and the suggestion of "if you were using them more actively you'd know this, this is a very common issue" is not at all weird. There are other ways they could have known this, but they didn't. "You haven't got the experience yet" is a much milder way of saying someone doesn't know how to use a tool properly than "you're ignorant".
- fwip 1y agoI think part of the issue is that it doesn't "feel" like the LLM is generating a URL, because that's not what a human would be doing. A human would be cut & pasting the URLs, or editing the code around them - not retyping them from scratch. Edit: I think I'm just regurgitating the article here.
- worldsayshi 1y agoThis is of course bad but: humans also makes (different) mistakes all the time. We could account for the risk of mistakes being introduced and make more tools that validate things for us. In a way LLM:s encourage us to do this by adding other vectors of chaos into our work. Like, why not have tools built into our environment that checks that links are not broken? With the right architecture we could have validations for most common mistakes without having the solution adding a bunch of tedious overhead.
- rossant 1y agoI agree, these kinds of stories should encourage us to setup more robust testing/backup/check strategies. Like you would absolutely have to do if you suddenly invited a bunch of inexperienced interns to edit your production code.
- rullelito 1y agoLLMs are turning into LLMs+hard-coded fixes for every imaginable problem.
- worldsayshi 1y agoWhy hard coded?
- exe34 1y ago> that checks that links are not broken? Can you spot the next problem introduced by this?
- cimi_ 1y agoYour point to not rely on good intentions and have systems in place to ensure quality is good - but your comparison to humans didn't go well with me. Very few humans fill in their task with made up crap then lie about it - I haven't met any in person. And if I did, I wouldn't want to work with them, even if they work 24/7. Obligatory disclaimer for future employers: I believe in AI, I use it, yada yada. The reason I'm commenting here is I don't believe we should normalise this standard of quality for production work.
- hshdhdhehd 1y agoWell using an LLM is like rolling dice. Logits are probabilities. It is a bullshit machine.
- dude250711 1y agoYeah, it read like "when running with scissors be careful out there". How about not running with scissors at all? Unless of course the management says "from now on you will be running with scissors and your performance will increase as a result".
- hansmayer 1y agoAnd if you stab yourself in the stomach ... you must have sucked at running with the scissors :)
- coldtea 1y ago>A few days later, just before deployment to production, I wanted to double check all 40 links. This was allowed to go to master without "git diff" after Codex was done?
- rossant 1y agoIt was a fairly big refactoring basically converting a working static HTML landing page into a Hugo website, splitting the HTML into multiple Hugo templates. I admit I was quite in a hurry and had to take shortcuts. I didn't have time to write automated tests and had to rely on manual tests for this single webpage. The diff was fairly big. It just didn't occur to me that the URLs would go through the LLMs and could be affected! Lesson learnt haha.
- exe34 1y agothis is why I'm terrified of large LLM slop changesets that I can't check side by side - but then that means I end up doing many small changes that are harder to describe in words than to just outright do.
- cimi_ 1y agoSpeaking of agents and tests, here's a fun one I had the other day: while refactoring a large code base I told the agent to do something precise to a specific module, refactor with the new change, then ensure the tests are passing. The test suite is slow and has many moving parts; the tests I asked it to run take ~5 minutes. The thing decided to kill the test run, then it made up another command it said was the 'tests' so when I looked at the agent console in the IDE everything seemed fine collapsed, i.e. 'Tests ran successfully'. Obviously the code changes also had a subtle bug that I only saw when pushing its refactoring to CI (and more waiting). At least there were tests to catch the problem.
- tuesdaynight 1y agoI think that it's something that model providers don't want to fix, because the amount of times that Claude Code just decided to delete tests that were not passing before I added a memory saying that it would need to ask for my permission to do that was staggering. It stopped happening after the memory, so I believe that it could be easily fixed by a system prompt.
- amelius 1y agoIn these cases I explicitly tell the llm to make as few changes as possible and I also run a diff. And then I reiterate with a new prompt if too many things changed.
- globular-toast 1y agoYou can always run a diff. But how good are people at reading diffs? Not very. It's the kind of thing you would probably want a computer to do. But now we've got the computer generating the diffs (which it's bad at) and humans verifying them (which they're also bad at).
- CaptainOfCoit 1y agoYeah, pick one for you to do, the other for the LLMs to do, ideally pick the one you're better at, otherwise 50/50 you'll actually become faster.
- Xss3 1y agoThis is a horror story about bad quality control practices, not the use of LLMs.
- __MatrixMan__ 1y agoI have a project that I've leaned heavily on LLM help for which I consider to embody good quality control practices. I had to get pretty creative to pull it off: spent a lot of time working on this sync system so that I can import sanitized production data into the project for every table it touches (there are maybe 500 of these) and then there's a bunch of hackery related to ensuring I can still get good test coverage even when some of these flows are partially specified (since adding new ones proceeds in several separate steps). If it was a project written by humans I'd say they were crazy for going so hard on testing. The quality control practices you need for safely letting an LLM run amok aren't just good. They're extreme.
- weinzierl 1y agoNot code, but I once pasted an event announcement and asked for just spelling and grammar check. LLM suggested a new version with minor tweak which I copy pasted back. Just before sending I noticed that it had moved the event date by one day. Luckily I caught it but it taught me that you never should blindly trust LLM output even with super simple tasks, no relevant context size, clear and simple one sentence prompt. LLM's do the most amazing things but they also sometimes screw up the simplest of tasks in the most unexpected ways.
- flowingfocus 1y agoA diff makes these kind of errors much easier to catch. Or maybe someone from XEROX has a better idea how to catch subtly altered numbers?
- mcpeepants 1y agoI verify all dates manually by memorizing their offset from the date of the signing of the Magna Carta
- nedrylandJP 1y agoHN is no place for chicanery.
- nonethewiser 1y ago>Not code, but I once pasted an event announcement and asked for just spelling and grammar check. LLM suggested a new version with minor tweak which I copy pasted back. Just before sending I noticed that it had moved the event date by one day. This is the kind of thing I immediately noticed about LLMs when I used them for the first time. Just anecdotally, I'd say it had this problem 30-40% of the time. As time has gone on, it has gotten so much better. But it still makes this kind of problem -- lets just say -- 5% of the time. The thing is, it's almost more dangerous to rarely make the problem. Because now people aren't constantly looking for it. You have no idea if it's not just randomly flipping terms or injecting garbage unless you actually validate it. The ideal of giving it an email to improve and then just scanning the result before firing it off is terrifying to me.
- grafmax 1y agoYeah this sort of thing is a huge time waster with LLMs.
- smougel 1y agoNot related to code... But when I use a LLM to perform a kind of copy/paste, I try to number the lines and ask it to generate a start_index and stop_index to perform the slice operation. Much less hallucinations and very cheap in token generation.
- yodsanklai 1y ago5 minutes ago, I asked Claude to add some debug statements in my code. It also silently changed a regex in the code. It was easily caught with the diff but can be harder to spot in larger changes.
- alzoid 1y agoI asked Claude to add a debug endpoint to my hardware device that just gave memory information. It wrote 2600 lines of C that gave information about every single aspect of the system. On the one hand kind of cool. It looked at the MQTT code and the update code, the platform (esp) and generated all kinds of code. It recommended platform settings that could enable more detailed information that checked out when I looked at the docs. I ran it and it worked. On the other hand, most of the code was just duplicated over and over again ex: 3 different endpoints that gave overlapping information. About half of the code generated fake data rather than actually do anything with the system. I rolled back and re-prompted and got something that looked good and worked. The LLMs are magic when they work well but they can throw a wrench into your system that will cost you more if you don't catch it. I also just had a 'senior' developer tell me that a feature in one of our platforms was deprecated. This was after I saw their code which did some wonky hacky like stuff to achieve something simple. I checked the docs and said feature (URL Rewriting) was obviously not deprecated. When I asked how they knew it was deprecated they said Chat GPT told them. So now they are fixing the fix chat gpt provided.
- troupo 1y ago> About half of the code generated fake data rather than actually do anything with the system. All the time // fake data. in production this would be real data ... proceeds to write sometimes hundreds of lines of code to provide fake data
- stuartjohnson12 1y ago"hey claude, please remove the fake data and use the real data" "sure thing, I'll add logic to check if the real data exists and only use the fake data as a fallback in case the real data doesn't exist"
- cpfohl 1y agoMy custom prompt instructs GPT to output changes to code as a diff/git-patch. I don’t use agents because it makes it hard to see what’s happening and I don’t trust them yet.
- ravila4 1y agoI’ve tried this approach when working in chat interfaces (as opposed to IDEs), but I often find it tricky to review diffs without the full context of the codebase. That said, your comment made me realize I could be using “git apply”more effectively to review LLM-generated changes directly in my repo. It’s actually a neat workflow!
- cpfohl 1y agoYep!! It’s fantastic
- mehdibl 1y agoErrors are normal and happen ofter. You need to focus on providing it ability to test the changes and fix errors. If you expect one shot you will get a lot of bad surprises.
- dkarl 1y agoI've had similar experience both in coding and in non-coding research questions. An LLM will do the first N right and fake its work on the rest. It even happens when asking an LLM to reformat a document, or asking it to do extra research to validate information. For example, before a recent trip to another city, I asked Gemini to prepare a list of brewery taprooms with certain information, and I discovered it had included locations that had been closed for years or had just been pop-ups. I asked it to add a link to the current hours for each taproom and remove locations that it couldn't verify were currently open, and it did this for about the first half of the list. For the last half, it made irrelevant changes to the entries and didn't remove any of the closed locations. Of course it enthusiastically reported that it had checked every location on the list.
- Romario77 1y agoLLMs are not good at "cycles" - when you have to go over a list and do the same action on each item. It's like it has ADHD and forgets or gets distracted in the middle. And the reason for that is that LLMs don't have memory and process the tokens, so as they keep going over the list the context becomes bigger with more irrelevant information and they can lose the reason they are doing what they are doing.
- polynomial 1y agoSo much for Difference and Repetition.
- steveklabnik 1y agoSurprised and a bit delighted to see a Deleuze reference on HN...
- dmoy 1y agoWhich is annoying because that is precisely the kind of boring rote programming tasks I want an LLM to do for me, to free up my time for more interesting problems
- 1y ago
- polynomial 1y agoEvals don't fix this.
- HardCodedBias 1y agoMaybe they don't fix it, but I suspect that they move us towards it occurring less often.
- FitchApps 1y agoReminds me when I asked Claude (through Windsurf) to create a S3 Lambda trigger to resize images (as soon as PNG image appears in S3, resize it). The code looked flawless and I deployed ..only to learn that I introduced a perpetual loop :) For every image resized, a new one would be created and resized. In 5 min, the trigger created hundreds of thousands of images ...what a joy was to clean that up in S3
- moomoo11 1y agoDo you write tests and do local testing?
- scottbez1 1y agoThe last point I think is most important: "very subtle and silently introduced mistakes" -- LLMs may be able to complete many tasks as well (or better) than humans, but that doesn't mean they complete them the same way, and that's critically important when considering failure modes. In particular, code review is one layer of the conventional swiss cheese model of preventing bugs, but code review becomes much less effective when suddenly the categories of errors to look out for change. When I review a PR with large code moves, it was historically relatively safe to assume that a block of code was moved as-is (sadly only an assumption because GitHub still doesn't have indicators of duplicated/moved code like Phabricator had 10 years ago...), so I can focus my attention on higher level concerns, like does the new API design make sense? But if an LLM did the refactor, I need to scrutinize every character that was touched in the block of code that was "moved" because, as the parent commenter points out, that "moved" code may have actually been ingested, summarized, then rewritten from scratch based on that summary. For this reason, I'm a big advocate of an "AI use" section in PR description templates; not because I care whether you used AI or not, but because some hints about where or how you used it will help me focus my efforts when reviewing your change, and tune the categories of errors I look out for.
- AIorNot 1y agoI think we need better code review tools in the age of LLMs - not just sticking another LLM to do a code review on top of the PR Needs to clearly handle the large diffs they produce - anyone have any ideas
- steveklabnik 1y agoI personally agree with you. I think that stacked diffs will be more important as a way of dealing with those larger diffs.
- ngruhn 1y agoI was about to write my own tool for this but then I discovered: git diff --color-moved=dimmed-zebra That shows a lot of code that was properly moved/copied in gray (even if it's an insertion). So gray stuff exactly matches something that was there before. Can also be enabled by default in the git config.
- BinaryIgor 1y ago"very subtle and silently introduced mistakes" - that's the biggest bottleneck I think; as long as it's true, we need to validate LLMs outputs; as long as we must validate LLMs outputs, our own biological brains are the ultimate bottleneck
- rapind 1y agoIncorrect data is a hard one to catch, even with automated tests (even in your tests, you're probably only checking the first link, if you're event doing that). Luckily I've grown a preference for statically typed, compiled, functional languages over the years, which eliminates an entire class of bugs AND hallucinations by catching them at compile time. Using a language that doesn't support null helps too. The quality of the code produced by agents (claude clode and codex) is insanely better than when I need to fix some legacy code written in a dynamic language. You'll sometimes catch the agent hallucinating and continuously banging it's head against the wall trying to get it's bad code to compile. It seems to get more desperate and may eventually figure out a way to insert some garbage to get it to compile or just delete a bunch of code and paper over it... but it's generally very obvious when it does this as long as you're reviewing. Combine this with git branches and a policy of frequent commits for greatest effect. You can probably get most of the way there with linters and automated tests with less strict dynamic languages, but... I don't see the point for new projects. I've even found Codex likes to occasionally make subtle improvements to code located in the same files but completely unrelated to the current task. It's like some form of AI OCD. Reviewing diffs is kind of essential, so using a foundation that reduces the size of those diffs and increases readability is IMO super important.
- thatfrenchguy 1y agoI truly wonder how much time we have before some spectacular failure will happen because a LLM was asked to rewrite a file with a bunch of constants in it in critical software and silently messed up or inverted them in a way that looks reasonable and works in your QA environment and then leads to a spectacular failure in the field.
- intrasight 1y agoAI coding and no automated testing is a bad combination.
- qnleigh 1y agoInteresting, I've seen similar looking behavior in other forms of data extraction. I took a picture of a bookshelf and asked it to list the books. It did well in the beginning but by the middle, it had started making up similar books that were not actually there.
- bryantnyc 1y ago"...very subtle and silently introduced mistakes are quite dangerous..." In my view these perfectly serve the purpose of encouraging you to keep burning tokens for immediate revenue as well as potentially using you to train their next model at your expense.