6 ms·
People’s acceptance of LLMs being flat wrong for most queries is so concerning. The fact that these technologies are being pushed onto the masses, onto healthc
by ado__dev 2y ago
People’s acceptance of LLMs being flat wrong for most queries is so concerning.
The fact that these technologies are being pushed onto the masses, onto healthcare, law, and other areas were accuracy is super important is going to get people killed.
But hey there’s a disclaimer that says “anything the LLM says may be wrong” so it’s all good.
I’m not a Luddite at all, I think LLMs are great for some use cases, but the way they’re being shoehorned into everything is a disaster waiting to happen.
- SOLAR_FIELDS 2y agoMaybe we can tone down the FUD a bit. Wikipedia is flat wrong sometimes. Google is flat wrong sometimes. LLM’s can be flat wrong sometimes. No different to trusting an LLM’s output to Google’s output. Good as a starting point but not something I’m going to base my medical and legal decisions on. I don’t see the zeitgeist of LLM’s being any different. I don’t see some trend of legal or medical professionals blindly trusting LLM output either, even if some very rare and cherry picked examples would want us to believe otherwise.
- r053bud 2y agoI must be living in a different world or using a different version of the model, but I seem to be getting garbage from OpenAI products, specifically ChatGPT 4o. The most recent example; I tried to recreate a similar scenario that Sal Khan did where he used photos/screenshare with ChatGPT to help his son learn geometry. I did the same thing with Chess. I started with having it walk me through a few chess puzzles. It straight up couldn’t figure out the solutions and frequently referenced coordinates that were well outside the bounds of a chess board (Knight to Q8, for example). At this point, it feels like I’m being gaslit by the community of AI obsessives who see LLMs as the second coming of Jesus. I get nothing but garbage from them. Sure, maybe it can occasionally write me a line of syntactically correct code but that’s it. It feels like this is being shoved down my throat and I’m criticized for ever expressing skepticism. I really don’t like how this discourse is progressing.
- csomar 2y agoThank god I am not the only one. I reverted back to GPT-4. 4o is complete garbage. It seems that it's geared toward always giving you an answer regardless whether it can determine an answer or not. It's hallucinating 80% of the time and with high confidence.
- j_bum 2y agoI’ve experienced excessive garbage production, but also excessive verbosity, even when it’s giving helpful responses. “Here’s the code I just wrote in the previous message!”
- delusional 2y agoIt's the Crypto problem again. Just ignore them. At some point their stupid technology will run out of money and they'll go away.
- disqard 2y agoI was about to write this as well. There are dozens of us! :)
- LightFog 2y agoDid you miss the bit where they are being marketed as ‘intelligence’ - non-techies gobble this stuff up.
- petesergeant 2y agoGell-Mann Amnesia effect is strong. LLM constantly hallucinates JS built-in functions and builds code with massive foot-guns, which I catch, but it's probably fine for pizza recipes ... right?
- ben_w 2y agoSo long as you don't mind glue on them, apparently. (I've never seen this category of mistake on ChatGPT, but it does have a really hard time understanding that I want metric not imperial).
- chrismcb 2y agoMaybe it is me being pedantic, but AIs don't hallucinate. They make stuff up (it is pretty much all they do) but to claim that is a hallucination attributes a trait to them that they don't have (shoot calling then AI does the same thing)
- Aloisius 2y agoThat ship has sailed. > hallucinate - When an artificial intelligence (= a computer system that has some of the qualities that the human brain has, such as the ability to produce language in a way that seems human) hallucinates, it produces false information - Cambridge dictionary > hallucination - computing : a plausible but false or misleading response generated by an artificial intelligence algorithm - Webster's dictionary
- petesergeant 2y agoI dunno, I think even in the expression of “make stuff up” there’s anthropomorphising
- zug_zug 2y agoI wouldn't draw a parallel between google's AI being bad and all LLMs.
- justapassenger 2y agoIt’s fundamentally same technology with same limitations. You can throw a lot of money at it to try to fix short comings, but they’ll be just bandaids. LLMs are extremely amazing autocomplete mechanisms. But that’s all they are at the core - there’s no intelligence involved.
- rowanG077 2y agoThat's a pretty rich statement considering they perform pretty good at tests which we have designed to measure intelligence. I don't see how you can say there is no intelligence.
- whythre 2y agoSounds like a problem with the tests being administered. There is a lot of woo around LLMs, a lot of people have a vested interest in hyping and selling AI; heck, even referring to LLMs as AI is a form of branding.
- rowanG077 2y agoSo turn it around then. When is an AI intelligent? Since when it performs at human standard in intelligence tests doesn't seem enough for you.
- blooalien 2y agohttps://arstechnica.com/science/2023/07/a-jargon-free-explanation-of-how-ai-large-language-models-work/ https://arstechnica.com/science/2023/07/a-jargon-free-explan... There's no intelligence involved in "artificial intelligence" as it currently stands. It's all marketing hype around a really fancy statistical completion engine. Intelligence would require thought and reasoning, which current "AI" does not do, no matter how convincingly it fakes it.
- A4ET8a8uTh0 2y agoI don't disagree. I saw it pushed by vendors at conference I attended recently ( regulatory compliance in that case ). Anyway, the presentation was pretty good, but the vendor knew that their target audience is on the conservative side when it comes to technology. The interesting part came up when they started moving into some details ( hallucinations and how to limit them ). They were selling 'AI job expert' ( compliance person focused on X ) that could be 'daisy chained' to limit it. I have to admit, it did not occur to me as an option and I feel like I should try it myself to see if it works as advertised.
- eastbound 2y agoCan anyone explain to me why it is not possible for LLMs to output the reference material together with the answer? In my ideal world, an AI assistant for professional situations would rather sound like an ideal HN answer which cites the reference, like “According to the law 12346-67, you are not allowed to add glue to food. But a researcher named John Doe in Arizona conducted an experiment with non-toxic glues with success. 67% of its participants didn’t die.” 3 sentences, 3 facts. Is it a feature that LLMs are keeping for future enterprise versions of their models, or is it entirely impossible?
- ipnon 2y agoperplexity.ai does exactly this very well. It's the best way to quickly find references for a problem you're thinking about.
- dragonwriter 2y ago> Can anyone explain to me why it is not possible for LLMs to output the reference material together with the answer? It can't do so in a particularly more reliable way (in a single pass) because every piece of input data potentially contributes to every response, and there is no deterministic way to "capture" which input work was meaningfully relevant to the output. You can ask an LLM to cite sources, which it will usually obediently do, but they may be incorrect or even completely invented-for-the-response sources. If you have access to an archive of the source data, you could do something like a semantically-aware search to try to attribute the answer to a work or multiple works in the training set, and if your setup is using RAG you can have the response (possibly bypassing the LLM, so this can be 100% reliable) identify any works that were consulted during the RAG step. You can't, of course, guarantee that the response was particularly focused on the cited source, while usually with a well-setup RAG setup the response should be informed by the retrieved documents, its possible that key aspects of the response are attributable to the general training of the model (and thus some other source or sources in the model) and not based on (and potentially even contradict) the doc(s) pulled inthe course of RAG.
- Gigachad 2y agoKind of like asking an artist what the source for their painting is. The source is a lifetime of viewing and practicing art, it’s not clear which pieces contributed which details to the final result.
- deleted 2y ago[deleted]