5 ms·
People have a fundamental misunderstanding of what LLMs are. I’ve had to explain this so many times even to engineers. People keep using it as Google. It is no
by LASR 2y ago
People have a fundamental misunderstanding of what LLMs are.
I’ve had to explain this so many times even to engineers. People keep using it as Google. It is not a mechanism to retrieve facts.
It’s rather a reasoning mechanism, that when used like a search engine generates text that looks like output from a retrieval of facts.
The article talks about OpenAI being unwilling to correct errors. But they just can’t. There just aren’t facts like birthdays for specific people discernible from the weights. Maybe for some.
So what would they correct? The best that can be done is like it is said in the article, apply filtering to refuse to answer.
- deleted 2y ago[deleted]
- Dalewyn 2y ago>People have a fundamental misunderstanding of what LLMs are. It's labeled "AI", aka "artificial intelligence". So clearly, it's intelligent enough to tell fact from fiction, right? It's intelligent enough to read and repeat facts, right? It's artificial intelligence, isn't it? I'm playing Devil's Advocate here, obviously. I'm aware how this actually works like you, but the way it's billed you can't expect the commons to understand any of this in any other way. There is a very specific concept of what AI is among the masses, regardless if "AI" is it or not. If you call something a duck, you can't blame someone for saying it's a duck.
- alpaca128 2y agoEven real intelligence struggles to tell facts from fiction, otherwise the world would look very different. That's yet another problem with AI, people often aren't sure what they can expect but many err on the side of "magical oracle".
- raincole 2y agoIs Gemini that unpopular? It seems that nobody in this thread is aware of its features. The best that can be done is just what Gemini does: a button to (smartly) compare the generated text against google search results.
- lnxg33k1 2y agoI don't know, it seems recently that we've detached what things are from what the people behind them say things are, so bing AI is an LLM, phind is an LLM, and they market themselves as search engine, like to see in details you say >> It is not a mechanism to retrieve facts Then bing AI powered by chatgpt shows on its site >> [Hello, this is Bing! I’m the new AI-powered chat mode of Microsoft Bing that] can help you quickly get information. if get is a synonym of retrieve and facts is a synonym of information, I'd argue that maybe it's not people misunderstanding what LLMs are but maybe someone explaining it wrong? Guess in the world of marketing words don't carry any value anymore, then I agree with you
- yunwal 2y agoBing is not just an LLM. It’s RAG, the LLM is just a layer on top of a typical search engine. The LLM is not “getting” any facts, it’s just synthesizing them in a human-readable format.
- tlrobinson 2y agoLLMs are also often used in the search component of RAG, by generating embeddings that are then indexed and searched. (I don’t know if that’s how Bing AI works)
- imiric 2y agoHow is LLM+RAG any different from a fine-tuned LLM without RAG? They're both trained on data, and have the capability to hallucinate.
- michaelt 2y agoTheoretically some people think RAG sounds more feasible to make factually accurate. After all, if you've trained an LLM on a masses of unchecked data you've scraped from the internet, your training data probably includes "Joe Biden is the president" and "Donald Trump is the president" and "Barrack Obama is the president" and "Emmanuel Macron est le président" and so on. It would be understandable if an LLM was confused about who the president was. These people think handing an LLM the contents of https://en.wikipedia.org/wiki/President_of_the_United_States https://en.wikipedia.org/wiki/President_of_the_United_States then asking who the president is sounds a lot more feasible. Personally I'm not so sure - I've never seen a RAG implementation that impressed me.
- emptyfile 2y ago[dead]
- trustno2 2y agoOpenAI openly and knowingly, and Microsoft even more, presents ChatGPT as a way to get facts.
- mschuster91 2y agoNow what I think would be interesting if one can use ChatGPT or whatever to distill any kind of document into a series of factoids. Say you give it the Wikipedia article about a famous person and it will extract stuff like "born: 1949-01-30" and associate it with the name of the person. Later on, a user asks an AI "when was Foo Bar born?", and the AI then looks up the factoid database and responds with the correct factoid, or an error message.
- karencarits 2y agoNote that a lot of the information on Wikipedia is available as a structured knowledge graph already: https://wikidata.org/wiki/Wikidata:Main_Page https://wikidata.org/wiki/Wikidata:Main_Page
- mschuster91 2y agoFor Wikipedia, yes, but I only used that as an example. There's tons of stuff one could build factoids from.
- input_sh 2y ago> So what would they correct? The best that can be done is like it is said in the article, apply filtering to refuse to answer. Stop providing service in the region in which your product is unable to comply with local regulations? You can't extract millions of euros from EU citizens with anything physical that goes against the EU law, but we like to pretend that because something's digital, it's totally fine if your product intentionally breaks the law, even if it's just for a couple of years until some random NGO sues you and courts react. I think that's nonsense. I'm not gonna say the EU needs its own equivalent of the Great Firewall, but there should be some cost of being intentionally non-compliant, as in fines, as in the same thing the EU already does to Facebooks and TikToks of the world.
- Mordisquitos 2y agoThe organic neural network inside my skull isn't a mechanism to retrieve facts either, nor are specific outputs that it may provide discernible from the weights. And yet, if it responds with a factual error to a given prompt, it can be trivially corrected so as to provide the right answer for future prompts.
- intelVISA 2y agoBold of you to assume 'Open' AI is accountable to anything other than printing more money f-for h-humanity of course.
- iLoveOncall 2y ago> The article talks about OpenAI being unwilling to correct errors. But they just can’t. You surely understand how this isn't a compelling argument at all, right? If you have a bug in your X-Ray software and it irradiates patients but you say you don't have the knowledge or resources to fix it, it doesn't suddenly become a way out of fixing it. This will be a slam dunk of a case. There are many obvious paths that OpenAI can follow, such as stopping business in the EU or preventing it from giving any fact about any person.
- michaelt 2y ago> The article talks about OpenAI being unwilling to correct errors. But they just can’t. There are actually several algorithms intended to allow fact editing in LLMs: https://github.com/zjunlp/EasyEdit?tab=readme-ov-file#current-implementation https://github.com/zjunlp/EasyEdit?tab=readme-ov-file#curren... They don't work perfectly (e.g. "Tim Cook is CEO of Apple" and "The CEO of Apple is Tim Cook" for some reason have to be edited separately) but they can deal with the most egregious cases. And, you know, maybe OpenAI can improve them further, given they're always going on about 'safety' as a reason their competitors should be regulated out of existence.
- Someone 2y ago> It’s rather a reasoning mechanism I wouldn’t call it reasoning. To me that implies using logic and being objective. It’s more wisdom/stupidity of the crowds, with the model designer deciding what crowd to use to create the model and then tweaking things to make the model look like it’s reasoning.