15 ms·
Ask HN: Share your "LLM screwed us over" stories?
Saw this today https://news.ycombinator.com/item?id=42575951 and thought that there might be more such cautionary tales. Please share your LLM horror stories for all of us to learn.
- byyoung3 2y agodoesn't seem like many people can come up with a specific and tangible example.
- Aachen 2y agoNot sure if this is what you're getting at, but the absence of a thing doesn't prove the opposite if you don't know that the right people are even on this site, much less reading this thread and deciding to put in the effort to comment after the thread existed for <1 hour Speaking for myself, I barely use LLMs and >50% of interactions are useless but I'm so wary that I doubt anything bad could happen. I also feel like I'm the target audience on HN, which someone who blindly trusts this magic black box (from their POV) might not be
- byyoung3 2y ago"the absence of a thing doesn't prove the opposite" - it definitely makes a strong case for the opposite
- talldayo 2y agoI think the conclusion in the thread is sound. If money means something to you, then don't ask AI for help spending it. Problem solved.
- fifilura 2y agoCould be extended to: If money means something to you don't ask a human or AI for help spending it.
- thisguy47 2y agoThe linked post is more a story of someone not understanding what they're deploying. If they had found a random blog post about spot instances, they likely would have made the same mistake. In this case, the LLM suggested a potentially reasonable approach and the author screwed themselves by not looking into what they were trading off for lower costs.
- SoftTalker 2y agoIt's also a story of AWS being so complicated in its various offerings and pricing that a person feels the need to ask an AI what's the best option because it's not obvious at all without a lot of study.
- steveBK123 2y agoIt’s pretty clear at this point AWS is purposely opaque.
- cornel_io 2y agoNo, they just ship their org chart. Which has grown incredibly, to the point where even internally nobody knows what's being shipped three branches over on the tree.
- daveguy 2y agoAgreed. You should always assume LLMs will give bad advice and verify information independently.
- deadbabe 2y agoFor the story above always ask the LLM to answer as if they are a Hackernews commenter before requesting advice.
- alecco 2y agoShare your "I used a tool and shot myself in the foot because I was lazy" stories.
- mmoustafa 2y agoIt's useful to see how other people shot themselves in the foot.
- __MatrixMan__ 2y agoI once used an LLM to generate some nix code to package a particular version of kustomize. It "worked", when I ran `kusomize version` in my shell, I got the version I asked for. I later realized that it wasn't packaging the desired version of kustomize, but was instead patching the build with a misleading version tag such that it had the desired version string but was in fact still the wrong version.
- aaroninsf 2y ago_Paperclip maximized!_
- bvanderveen 2y agoI'm surprised at how even some of the smartest people in my life take the output of LLMs at face value. LLMs are great for "plan a 5 year old's birthday party, dinosaur theme", "design a work-out routine to give me a big butt", or even rubber-ducking through a problem. But for anything where the numbers, dates, and facts matter, why even bother?
- jasfi 2y agoThey can sometimes give useful advice. But remember to see them as limited, and that they can make mistakes or leave things out.
- daveguy 2y agoI wouldn't trust an LLM for anything, especially exercise or anything close to medical advice.
- steveBK123 2y agoPeople enjoy being told what to do in some cases / planning is not a common trait
- vrx-meta 2y agoI have noticed I sometimes prompt in such a way that it outputs more or less what I already want to hear. I seek validation from LLMs. I wonder what could go wrong here.
- danielbln 2y agoYou're basically leading the witness. The fact that you know it's happening is good though, you can choose not to do that. Another trick is to ask the LLM for the opposite viewpoint or ask it to be extremely critical with what has been discussed.
- dismalaf 2y agoMy test for LLMs (mainly because I love cooking): "Give me a recipe for beef bourguignon" Half the time I get a chicken recipe...
- olalonde 2y agoNot me but Craig Wright aka Faketoshi referenced court cases hallucinated by a LLM in his appeal. https://cointelegraph.com/news/court-rejects-craig-wright-appeal-bitcoin-creator-case https://cointelegraph.com/news/court-rejects-craig-wright-ap...
- woolion 2y agoI could never convince a senior engineer that even very basic legal financial questions get very wrong answer by ChatGPT. "I don't even need a lawyer's services anymore", he said. This is in Belgium, so I don't think it's even reasonable to assume that it would be very accurately localized.
- lunarcave 2y agoPerhaps the story that doesn't get told more often is how LLMs are changing how humans operate en masse. When ChatGPT came out, I was increasingly outsourcing my thinking to LMS. It took me a few months to figure out that that's actually harming me - I've lost my ability to think through things a little bit. The same is true for Coding Assistants; sometimes I disable the in-editor coding suggestions, when I find that my coding has atrophied. I don't think this is necessarily a bad thing, as long as LMs are ubiquitous and they proliferate throughout society and are extremely reliable and accessible. But they are not there today.
- ok123456 2y agoPeople said the same thing about IntelliSense. The LLM should free you from thinking about relatively unimportant details of programming on the small and allow you to experiment with how things fit together. If you use the LLM to churn out garbage as is, it's essentially the same as performing random auto-completes with IntelliSense and just going with whatever pops out. Yes, that will hurt you professionally.
- __loam 2y agoCars free me from the drudgery of walking to places but I still ride my bike to stay fit. You are an energy optimization machine and lots of your biology is use it or lose it.
- harrisonjackson 2y agoI look at it more like a lever or pulley. I'm still going to get a good workout exerting the maximum that I can. But wow can I lift a lot more.
- sylos 2y agoCars are actually a good analogy. They seem helpful in a lot of ways, and they can be, but ultimately they're harmful to your health, other's health, environment, and the design and ability to construct places and locations free of drudgery.
- bflesch 2y agoLLMs = ad-free version of google That's why people adopted it. Google got worse and worse, now the gap is filled with LLMs. LLMs have replaced google, and that's awesome. LLMs won't cook lunch or fold our laundry, and until a better technology comes around which can actually do that all promises around "AI" should be seen as grifting.
- richardwhiuk 2y agoThe ads will come and they'll need to be an order of magnitude worse to make llms profitable.
- oezi 2y agoLLMs know a magnitude more about you and your needs. This will make targeting users very valuable. The big question is who will be first to get it done in an unobstrusive way (subtlely integrated into text with links).
- HWR_14 2y agoI suppose the only solution is to communicate with LLMs only through TOR and pay for them only via crypto?
- kristopolous 2y agoThat's the real AI future, not some atomic war with robots. Manipulative AI bowing to the dollars of advertisers, responding to something overheard by a smart speaker 20 minutes ago by packaging an interstitial ad for coca-cola and delivering it stealthily in-experience to your child in their bedroom on a VR headset, straight to their cornea. The ways for circumventing the influence are currently being dismantled and these AI RTB ad systems are well funded and being built. AI news feed will echo the message and the advertiser will be there again in their AI internet search. We will cede the agencies of reality to machines responding to those wishing to reshape reality in their own interest as more of the human experience gets commodified, packaged, and traded as a security to investors looking to extract value from being alive.
- 2y ago
- troilboil 2y agoNot mine but a client of mine. Consultants sold them a tool that didn't exist because the LLM hallucinated and told their salesperson it did. Not sure that's really the LLM's fault, but pretty funny.
- datadrivenangel 2y agoSo the consultants were told by an LLM that they sold a tool which they actually didn't? Was this chatgpt or something custom?
- mtmail 2y agoChatGPT claims our service has a feature which we don't have (for example tracking people based on their phone number). Users register a free account, then complain to us. The first email is often vague "It doesn't work" without details. Slightly worse is users who go ahead and make a purchase, then complain, then demand a refund. We had to add a warning on the account registration page.
- LeoPanthera 2y agoI'm currently shopping for a new car, and while asking questions at a dealer (not Tesla), they revealed that the sales guys use ChatGPT to look up information about the car because it's quicker than trying to find things in their own database. I did not buy that car.
- woolion 2y agoI've tried LLMs for a few exploratory programming projects. It kinda feels magical the first time you import a dependency you don't know and you get the LLM output what you want to do without you even having the time to think about it. However, I also think that for any minute I've gained with it I've lost at least one because of hallucinated solutions. Even for fairly popular things (Terraform+AWS) I continuously got plausible-looking answers. After reading carefully the docs, the use case was not supported at all, so I just went with the 30 seconds (inefficient) solution I had thought of from the start. But I lost more than one hour. Same story with the Ren'py framework. The issue is that the docs are far from covering everything, and Google sucks, sometimes giving you a decade-old answer to a problem that has a fairly good answer in more recent versions. So it's really difficult to decide how to most efficiently look for an answer between search and LLM. Both can be a stupid waste of time.
- ldehaan 2y ago[dead]
- 0xDEAFBEAD 2y agoI find it interesting how LLM errors can be so subtle. The next-token prediction method rewards superficial plausibility, so mistakes can be hard to catch.
- starchild3001 2y agoGuys, it's a major 21st century skill to learn how to use LLMs. In fact, it's probably the biggest skill anyone can develop today. So please be a responsible driver, learn how to use LLMs. Here's one way to get the most mileage out of them: 1) Track the best and brightest LLMs via leaderboards (e.g. https://lmarena.ai/ https://lmarena.ai/, https://livebench.ai/#/ https://livebench.ai/#/ ...). Don't use any s**t LLMs. 2) Make it a habit to feed in whole documents and ask questions about them vs asking them to retrieve from memory. 3) Ask the same question to the top ~3 LLMs in parallel (e.g. top of line Gemini, OpenAI and Claude models) 4) Do comparisons between results. Pick best. Iterate on the the prompt, question and inputs as required. 5) Validate any key factual information via Google or another search engine before accepting it as a fact. I'm literally paying for all three top AIs. It's been working great for my compute and information needs. Even if one hallucinates, it's rare that all three hallucinate the same thing at the same time. The quality has been fantastic, and intelligence multiplication is supreme.
- jpcookie 2y agoGive me a real example of something in computer science they can't do? I'm interested since Chatgpt is better than any professor I've had at any level in my educational career.