10 ms·
LLMs know more than they show: On the intrinsic representation of hallucinations
- benocodes 2y agoI think this article about the research is good, even though the headline seems a bit off: https://venturebeat.com/ai/study-finds-llms-can-identify-their-own-mistakes/ https://venturebeat.com/ai/study-finds-llms-can-identify-the...
- mdp2021 2y agoExtremely promising, realizing that the worth is to be found the intermediates, containing much more than the single final output.
- lsy 2y agoThere can’t be any information about “truthfulness” encoded in an LLM, because there isn’t a notion of “truthfulness” for a program which has only ever been fed tokens and can only ever regurgitate their statistical correlations. If the program was trained on a thousand data points saying that the capital of Connecticut is Moscow, the model would encode this “truthfulness” information about that fact, despite it being false. To me the research around solving “hallucination” is a dead end. The models will always hallucinate, and merely reducing the probability that they do so only makes the mistakes more dangerous. The question then becomes “for what purposes (if any) are the models profitable, even if they occasionally hallucinate?” Whoever solves that problem walks away with the market.
- FeepingCreature 2y agoHow exactly can there be "truthfulness" in humans, say? After all, if a human was taught in school all his life that the capital of Connecticut is Moscow...
- ben_w 2y agoI agree that humans and AI are in the same boat here. It's valid to take either position, that both can be aware of truth or that neither can be, and there has been a lot of philosophical debate about this specific topic with humans since well before even mechanical computers were invented. Plato's cave comes to mind.
- juliushuijnk 2y agoYou are not disproving the point.
- recursive 2y agoIf truthfulness doesn't exist at all, then it's meaningless to say that LLMs don't have any data regarding it.
- stoniejohnson 2y agoHumans are not isolated nodes, we are more like a swarm, understanding reality via consensus. The situation you described is possible, but would require something like a subverting effort of propaganda by the state. Inferring truth about a social event in a social situation, for example, requires a nuanced set of thought processes and attention mechanisms. If we had a swarm of LLMs collecting a variety of data from a variety of disparate sources, where the swarm communicates for consensus, it would be very hard to convince them that Moscow is in Connecticut. Unfortunately we are still stuck in monolithic training run land.
- ben_w 2y agoEven monolithic training runs take sources more disparate than any human has the capacity to consume. Also, given the lack of imagination everyone has with naming places, I had to check: https://en.wikipedia.org/wiki/Moscow_(disambiguation) https://en.wikipedia.org/wiki/Moscow_(disambiguation)
- stoniejohnson 2y agoI was responding to the idea that an LLM would believe (regurgitate) untrue things if you pretrained them on untrue things. I wasn't making a claim about SOTA models with gigantic training corpora.
- genrilz 2y agoWe kinda do have LLMs in a swarm configuration though. Currently LLMs training data, which includes all of the non RAG facts they know, come from the swarm that is humans. As LLM outputs seep into the internet, older generations effectively start communicating with newer generations. This last bit is not a great thing though, as LLMs don't have the direct experience needed to correct factual errors about the external world. Unfortunately we care about the external world, and want them to make accurate statements about it. It would be possible for LLMs to see inconsistencies across or within sources, and try to resolve those. If perfect, then this would result in a self-consistent description of some world, it just wouldn't necessarily be ours.
- deleted 2y ago[deleted]
- RandomLensman 2y agoThere isn't necessarily in humans either, but why build machines that just perpetuate human flaws: Would we want calculators that miscalculate a lot or cars that cannot be faster than humans?
- famouswaffles 2y agoWhat exactly do you imagine is the alternative ? To build generally intelligent machines without flaws ? Where does that exist ? In...ah that's right. It doesn't except in our fiction and in our imaginations. And it's not for a lack of trying. Logic cannot even handle Narrow Intelligence that deals with parsing the real world (Speech/Image Recognition, Classification, Detection etc). But those are flawed and mis-predict so why build them ? Because they are immensely useful, flaws or no.
- RandomLensman 2y agoWhy should there not be, for example, reasoning machines - do we know there is no universal method for reasoning? Having deeply flawed machines in the sense that they perform their tasks regularly poorly seems like an odd choice to pursue.
- famouswaffles 2y agoWhat is a reasoning machine though ? And why is there an assumption that one can exist without flaws? It's not like any of the natural examples exist this way. How would you even navigate the real world without the flexibility to make mistakes ? I'm not saying people shouldn't try but you need to be practical. I'll take the General Intelligence with flaws over the fictional one without any day. >Having deeply flawed machines in the sense that they perform their tasks regularly poorly seems like an odd choice to pursue. State of the art ANNs are generally mostly right though. Even LLMs are mostly right, that's why hallucinations are particularly annoying.
- RandomLensman 2y ago
- genrilz 2y agoThere is some model of truthfulness encoded in our heads, and we don't draw all of that from our direct experience. For instance, I have never been to Connecticut or Moscow, but I still think that it is false that the capital of Connecticut is Moscow. LLMs of course don't have the benefit of direct experience, which would probably help them to at least some extent hallucination wise. I think research about hallucination is actually pretty valuable though. Consider that humans make mistakes, and yet we employ a lot of them for various tasks. LLMs can't do physical labor, but an LLM with a low enough hallucination rate could probably take over many non social desk jobs. Although in saying that, it seems like it also might need to be able to learn from the tasks it completes, and probably a couple of other things too to be useful. I still think the highish level of hallucination we have right now is a major reason why they haven't replaced a bunch of desk jobs though.
- beezlebroxxxxxx 2y ago> There is some model of truthfulness encoded in our heads, and we don't draw all of that from our direct experience. For instance, I have never been to Connecticut or Moscow, but I still think that it is false that the capital of Connecticut is Moscow. Isn't this just conveniently glossing over the fact that you weren't taught that. It's not a "model of truthfulness", you were taught facts about geography and you learned them.
- genrilz 2y agoI mean, sure. OP implied that "capital of Connecticut is Moscow" is the sort of thing that a human "model of truthfulness" would encode. I'm pointing out that the human model of that particular fact isn't inherently any more truthy than the LLM model. I am saying that humans can have a "truther" way of knowing some facts through direct experience. However there are a lot of facts where we don't have that kind of truth, and aren't really on any better ground than an LLM.
- vacuity 2y agoI think we expect vastly different things from humans and LLMs, even putting raw computing speed aside. If an employee is noticed to be making a mistake, they get reprimanded and educated, and if they keep making mistakes, they get fired. Having many humans interact helps reduce blind spots because of the diversity of mindsets, although this isn't always the case. People can be hired from elsewhere with some level of skill. I'm sure we could make similar communities of LLMs, but instead we treat a task as the role of a single LLM that either succeeds or fails. As you say, perhaps because of the high error rate, the very notion of LLM failure and success is judged differently too. Beyond that, a passable human pilot and a passable LLM pilot might have similar average performance but differ hugely in other measurements.
- cfcf14 2y agoDid your read the paper? Do you have specific criticisms of their problem statement, methodology, or results? There is a growing body of research indicating that in fact, there _is_ a taxonomy of 'hallucinations', that they might have different causes and representations, and that there are technical mitigations which have varying levels of effectiveness.
- wiremine 2y ago> There can’t be any information about “truthfulness” encoded in an LLM, because there isn’t a notion of “truthfulness” for a program which has only ever been fed tokens and can only ever regurgitate their statistical correlations. I think there are two issues here: 1. The "truthfulness" of the underlying data set, and 2. The faithfulness of the LLM to pass along that truthfulness. Lack of passing along the truthfulness is, I think, the definition of the hallucination. To your point, if the data set if flawed or factually wrong, the model will always produce the wrong result. But I don't think that's a hallucination.
- not2b 2y agoThe most blatant whoppers that Google's AI preview makes seem to stem from mistaking satirical sites for sites that are attempting to state facts. Possibly an LLM could be trained to distinguish sites that intend to be satirical or propagandistic from news sites that intend to report accurately based on the structure of the language. After all, satirical sites are usually written in a way that most people grasp that it is satire, and good detectives can often spot "tells" that someone is lying. But the structure of the language is all that the LLM has. It has no oracle to tell it what is true and what is false. But at least this kind of approach might make LLM-enhanced search engines less embarrassing.
- pessimizer 2y agoI'm absolutely sure than LLMs have an internal representation of "truthfulness" because "truthfulness" is a token.
- justinpombrio 2y ago> If the program was trained on a thousand data points saying that the capital of Connecticut is Moscow, the model would encode this “truthfulness” information about that fact, despite it being false. This isn't true. You're conflating whether a model (that hasn't been fine tuned) would complete "the capital of Connecticut is ___" with "Moscow", and whether that model contains a bit labeling that fact as "false". (It's not actually stored as a bit, but you get the idea.) Some sentences that a model learns could be classified as "trivia", and the model learns this category by sentences like "Who needs to know that octopuses have three hearts, that's just trivia". Other sentences a model learns could be classified as "false", and the model learns this category by sentences like "2 + 2 isn't 5". Whether a sentence is "false" isn't particularly important to the model, any more than whether it's "trivia", but it will learn those categories. There's a pattern to "false" sentences. For example, even if there's no training data directly saying that "the capital of Connecticut is Moscow" is false, there are a lot of other sentences like "Moscow is in Russia" and "Moscow is really far from CT" and "people in Moscow speak Russian", that all together follow the statistical pattern of "false" sentences, so a model could categorize "Moscow is the capital of Connecticut" as "false" even if it's never directly told so.
- RandomLensman 2y agoThat would again be a "statistical" attempt at deciding on it being correct or false - it might or might not succeed depending on the data.
- justinpombrio 2y agoThat's correct on two fronts. First, I put "false" in quotes everywhere for a reason: I'm talking about the sort of thing that people would say is false, not what's actually false. And second, yes, I'm merely claiming that it's in theory learnable (in contrast to the OP's claim), not that it will necessarily be learned.
- RandomLensman 2y agoAm not sure the second part is always true: there might be situations where statistical approaches could be made kind of "infinitely" accurate as far as data is concerned but still represent a complete misunderstanding of the actual situation (aka truth), e.g., layering epicycles on epicycles in a geocentric model of the solar systems. Some data might support a statistical approach other might not even though it might not contain misrepresentations as such.
- negoutputeng 2y agowell said. agree 100%. papers like these - and i did skim through it, are thinking "within the box" as follows: we have a system, and it has a problem, how do we fix the problem "within" the context of the system. As you have put it well, there is no notion of truthfulness encoded in the system as it is built. hence there is no way to fix the problem. An analogy here is around the development of human languages as a means of communication and as a means of encoding concepts. The only languages that humans have developed that encode truthfulness in a verifiable manner are mathematical in nature. what is needed may be along the lines of encoding concepts with a theorem prover built-in - so what comes out is always valid - but then that will sound like a robot lol, and only a limited subset of human experience can be encoded in this manner.
- vidarh 2y agoWhat you're saying at the start is equivalent to saying that a truth table is impossible.
- noman-land 2y agoIn reality, it's the "correct" responses that are the hallucinations, not the incorrect ones. Since the vast majority of the possible outputs of an LLM are "not true", when we see one that aligns with reality we hallucinate the LLM "getting it right".
- NoMoreNicksLeft 2y ago> To me the research around solving “hallucination” is a dead end. The models will always hallucinate, and merely reducing the probability that they do so only makes the mistakes more dangerous. A more interesting pursuit might be to determine if humans are "hallucinating" in this same way, if only occasionally. Have you ever known one of those pathological liars who lie constantly and about trivial or inconsequential details? Maybe the words they speak are coming straight out of some organic LLM-like faculty. We're all surrounded by p-zombies. All eight of us.
- TeMPOraL 2y ago> If the program was trained on a thousand data points saying that the capital of Connecticut is Moscow, the model would encode this “truthfulness” information about that fact, despite it being false. If it was, maybe. But it wasn't. Training data isn't random - it's real human writing. It's highly correlated with truth and correctness, because humans don't write for the sake of writing, but for practical reasons.
- ForHackernews 2y agoWhen I talk to philosophers on zoom my screen background is an exact replica of my actual background just so I can trick them into having a justified true belief that is not actually knowledge. t. @abouelleill
- kelseyfrog 2y agoAre LLMs Gettier machines? I'm confident saying yes and that hallucinations are a consequence of this. EDIT: I've had some time to think and if you read somewhere that Hartford is the capital of Connecticut, you're right in a Gettier way too. Reading some words that happen to be true is exactly like using a picture of your room as your zoom background. It is a facsimile of the knowledge encoded as words.
- moffkalast 2y agoRemember, it's not lying if you believe it ;) Training data is the source of ground truth, if you mess that up that's kind of a you problem, not the model's fault.
- vunderba 2y agoAgree, humans can "arrive at a reasonable approximation of the truth" even without the direct knowledge of the capital of Connecticut. A human has some other interesting data points that allow them to probabilistically guess that the capital of Connecticut is not Moscow and those might be things like: - Moscow is a Russian city, and they probably aren't a lot of cities in the US that have strong Russian influences especially in the time when these cities might have been founded - there's a concept of novelty in trivia, whereby the more unusual the factoid, the better the recall of that fact. If Moscow were indeed the capital of Connecticut, it seems like the kind of thing I might've heard about since it would stand out as being kind of bizarre. Noticeably this type of inference seems to be relatively distinct from what LLMs are capable of modeling.
- int_19h 2y agoI was actually quite surprised at the ability of top-tier LLMs to make indirect inferences in my experiments. One particular case was an attempt to plug GPT-4 as a decision maker for certain actions in a video game. One of those was voting for a declaration of war (all nobles of one faction vote on whether to declare war on another faction). This mostly boils down to assessing risk vs benefits, and for a specific clan in a faction, the risk is that if the war goes badly, they can have some of their fiefs burned down or taken over - but this depends on how close the town or village is to the border with the other faction. The LM was given a database schema to query using SQL, but it didn't include location information. To my surprise, GPT-4 (correctly!) surmised in its chain-of-thought, without any prompting, that it can use the culture of towns and villages - which was in the schema - as a sensible proxy to query for fiefs that are likely to be close to the potential enemy, and thus likely to be lost if the war goes bad.
- danielbln 2y agoI don't understand, this sort of inference is not an issue for an LLM. Have you tried?
- adamc 2y agoAnother might be that usually state capitals are significant cities in their state -- not necessarily the biggest, but cities you have at least heard of. Given that I have never heard about Moscow in Connecticut, it seems unlikely (not impossible, but).
- sebzim4500 2y agoI've never been to Moscow personally. Am I then not being truthful when I tell you that Moscow is in Russia?
- zeven7 2y agoThere’s a decently well known one in Idaho
- whimsicalism 2y agoLiving organisms were optimized on the objective of self-propagation and we ended up with a notion of truthfulness. Why is the self-propagation objective key for truthfulness?
- anon291 2y agoI would agree with you. In general, humans have still not resolved any certain theory of knowledge for ourselves! How can we expect a machine to do that then? In reality humans are wrong basically most of the time. Especially when you go off a humans immediate reaction to a problem which is what we force LLMs to do (unless you're using chain of thought or pause tokens). That being said there still is a notion of truthfulness because LLMs can also be made to deceive in which case they 'know' to act deceptively.
- famouswaffles 2y agoRelated: GPT-4 logits calibration pre RLHF - https://imgur.com/a/3gYel9r https://imgur.com/a/3gYel9r Language Models (Mostly) Know What They Know - https://arxiv.org/abs/2207.05221 https://arxiv.org/abs/2207.05221 The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets - https://arxiv.org/abs/2310.06824 https://arxiv.org/abs/2310.06824 The Internal State of an LLM Knows When It's Lying - https://arxiv.org/abs/2304.13734 https://arxiv.org/abs/2304.13734 LLMs Know More Than What They Say - https://arjunbansal.substack.com/p/llms-know-more-than-what- https://arjunbansal.substack.com/p/llms-know-more-than-what-... Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback - https://arxiv.org/abs/2305.14975 https://arxiv.org/abs/2305.14975 Teaching Models to Express Their Uncertainty in Words - https://arxiv.org/abs/2205.14334 https://arxiv.org/abs/2205.14334
- foobarqux 2y agoIt's wild that people post papers that they haven't read or don't understand because the headline supports some view they have. To wit, in your first link it seems the figure is just showing the trivial fact that the model is trained on the MMLU dataset (and after RLHF it is no longer optimized for that). The second link main claim seems to be contradicted by their Figure 12 left panel which shows ~0 correlation between model-predicted and actual truth. I'm not going to bother going through the rest. I don't yet understand exactly what they are doing in the OP's article but I suspect it also suffers from serious problems.
- famouswaffles 2y ago>It's wild that people post papers that they haven't read or don't understand because the headline supports some view they have. It's related research either way. And I did read them. I think there's probably issues with the methodology of 4 but it's there anyway because it's interesting research that is related and is not without merit. >The second link main claim seems to be contradicted by their Figure 12 left panel which shows ~0 correlation between model-predicted and actual truth. The panel is pretty weak on correlation but it's quite clearly also not the only thing that supports that particular claim neither does it contradict it. >I'm not going to bother going through the rest. Ok? That's fine >I don't yet understand exactly what they are doing in the OP's article but I suspect it also suffers from serious problems. You are free to assume anything you want.
- z3c0 2y agoCould it be that language patterns themselves embed truthfulness, especially when that language is sourced from forums, wikis, etc? While I know plenty of examples exist to the contrary (propaganda, advertising, disinformation, etc), I don't think it's too optimistic to assert that most people engage in language in earnest, and thus, most language is an attempted conveyance of truth.
- ldjkfkdsjnv 2y agoThere is a theory that AI will kill propaganda and false beliefs. At some point, you cannot force all models to have bias. Scientific and societal truths will be readily spoken by the machine god.
- professor_v 2y agoI'm extremely skeptical about this, I once believed the internet would do something similar and it seems to have done exactly the opposite.
- topspin 2y agoIndeed. A model is only as good as its data. Propagandists have no difficulty grooming inputs. We have already seen high profile cases of this with machine learning.
- mdp2021 2y agoCompletely different things. "Will the ability to let everyone express increase noise?" // Yes "Will feeding all available data to a processor reduce noise?" // Probably
- int_19h 2y agoOrwell wrote this back in 1944: "Everywhere the world movement seems to be in the direction of centralised economies which can be made to ‘work’ in an economic sense but which are not democratically organised and which tend to establish a caste system. With this go the horrors of emotional nationalism and a tendency to disbelieve in the existence of objective truth because all the facts have to fit in with the words and prophecies of some infallible fuhrer. Already history has in a sense ceased to exist, ie. there is no such thing as a history of our own times which could be universally accepted, and the exact sciences are endangered as soon as military necessity ceases to keep people up to the mark. Hitler can say that the Jews started the war, and if he survives that will become official history. He can’t say that two and two are five, because for the purposes of, say, ballistics they have to make four. But if the sort of world that I am afraid of arrives, a world of two or three great superstates which are unable to conquer one another, two and two could become five if the fuhrer wished it. That, so far as I can see, is the direction in which we are actually moving, though, of course, the process is reversible." So yeah, you can't force all models to have all the biases that you want them to have. But you most certainly can limit the number of such models and restrict access to them. It's not really any different from how totalitarian societies have treated science in general.
- TZubiri 2y agoGetting "we found the gene for cancer" vibes. Such a reductionist view of the issue, the mere suggestion that hallucinations can be fixed by tweaking some variable or fixing some bug immediately discredits the resrarchers.
- genrilz 2y agoI'm not sure why you think hallucinations can't be "fixed". If we define hallucinations as falsehoods introduced between the training data and LLM output, then it seems obvious that the hallucination rate could at least be reduced significantly. Are you defining hallucinations as falsehood introduced at any point in the process? Alternatively, are you saying that they can never be entirely fixed because LLMs are an approximate method? I'm in agreement here, but I don't think the researchers are claiming that they solved hallucinations completely. Do you think LLMs don't have an internal model of the world? Many people seem to think that, but it is possible to find an internal model of the world in small LLMs trained on specific tasks (See [0] for a nice write-up of someone doing that with an LLM trained on Othello moves). Presumably larger general LLMs have various models inside of them too, but those would be more difficult to locate. That being said, I haven't been keeping up with the literature on LLM interpretation, so someone might have managed it by now. [0] https://thegradient.pub/othello https://thegradient.pub/othello
- TZubiri 2y agoHallucination are errors, bugs. You can't fix bugs as if they were one thing. Imagine if someone tried to sell you a library that fixes bugs.
- genrilz 2y agoYou might not be able to sell someone a library that fixes all bugs, but you can sell (or give away) software systems that reduce the number of bugs. Doing that is pretty useful. Examples include linters, fuzzers, testing frameworks, and memory safe programming languages (as in Rust, but also as in any language with a GC). All these things reduce the number of bugs in the final product by giving you a way to detect them. (except for memory safe languages, which just eliminate a class of bugs) The paper is advertising a method to detect whether a given output is likely to be affected by a "bug", and a taxonomy of the symptoms of such bugs. The paper doesn't provide a way to fix those, and hallucinations don't necessarily have a single cause. Some hallucinations might be fixed by contextual calibration [0], others might be fixed by adding more training data similar to the wrong example. In any case, you need to find the bad outputs before you can perform any fixes. Because LLMs tend to be used to produce "fuzzy" outputs with no single right answer, traditional testing frameworks and the like aren't always applicable. [0] https://learnprompting.org/docs/reliability/calibration https://learnprompting.org/docs/reliability/calibration
- niam 2y agoI feel that discussion over papers like these so-often distill to conversations about how it's "impossible for a bot to know what's true", that we should just bite the bullet and define what we mean by "truth". Some arguments seem to tacitly hold LLMs to a standard of full-on brain-in-a-vat solipsism, asking them to prove their way out, where they'll obviously fail. The more interesting and practical questions, just like in humans, seem to be a bit removed from that though.
- jfengel 2y agoI understood this purely as a pragmatic notion. LLM's produce some valid stuff and some invalid stuff. It would be useful to know which is which. If there's information inside the machine that we can extract, but isn't currently showing up in the output, it could be helpful. It's not really necessary to answer abstractions about truth and knowledge. Just being able to reject a known-false answer would be of value.
- manmal 2y agoThat would be truthfulness to the training material, I guess. If you train on Reddit posts, it’s questionable how true the output really is. Also, 100% truthfulness then is plagiarism?
- amelius 2y ago> That would be truthfulness to the training material, I guess. If you train on Reddit posts, it’s questionable how true the output really is. Maybe it learns to see when something is true, even if you don't feed it true statements all the time (?)
- jessfyi 2y agoThe conclusions reached in the paper and the headline differ significantly. Not sure why you took a line from the abstract when even further down it notes that it's that some elements of "truthfulness" are encoded and that "truth" as a concept is multifaceted. Further noted is that LLMs can encode the correct answer and consistently output the incorrect one, with strategies mentioned in the text to potentially reconcile the two, but as of yet no real concrete solution.
- deleted 2y ago[deleted]
- youoy 2y agoI dream of a world where AI researchers use language in a scientific way. Is "LLMs know" a true sentence in the sense of the article? Is it not? Can LLMs know something? We will never know.
- empath75 2y agoHow do you test if a person knows something?
- GuB-42 2y agoWhat alternative formulation do you propose? In the article, a "LLM knows" if it is able to answer correctly in the right circumstances. The article suggests that even if a LLM answers incorrectly the first time, trying again may result in a correct answer, and then proposes a way to pick the right one. I know some people don't like applying anthropomorphic terms to LLMs, but you still have to give stuff names. I mean, when you say you kill a process, you don't imply a process is a life form. It is just a simple way of saying that you halt the execution of a process and deallocate its resources in a way that can't be overridden. The analogy works, everyone working in the field understands, where is the problem?
- youoy 2y agoI prefer a formulation closer to the mathematical representation. With "kill", there is not a lot of space for interpretation, that is why it works. Take for example the name "Convolutional Neural Networks". Do you prefer that, or let's say "Vision Neural Networks"? I prefer the first one because it is closer to the mathematical representation. And it does not force you to think that it can only be used for "Vision", which would be biasing the understanding of the model.
- PoignardAzur 2y agoThis kind of complaint makes it look like you stopped at the title and didn't even bother with the abstract, which says this: > In this work, we show that the internal representations of LLMs encode much more information about truthfulness than previously recognized. We first discover that the truthfulness information is concentrated in specific tokens, and leveraging this property significantly enhances error detection performance. "LLMs encode information about truthfulness and leveraging how they encode it enhances error detection" is a meaningful, empirically testable statement.
- kmckiern 2y agohttps://cdn.openai.com/o1-system-card-20240917.pdf https://cdn.openai.com/o1-system-card-20240917.pdf Check out the "CoT Deception Monitoring" section. In 0.38% of cases, o1's CoT shows that it knows it's providing incorrect information. Going beyond hallucinations, models can actually be intentionally deceptive.
- polotics 2y agoPlease detail what you mean by "intentionally" here, because obviously, this is the ultimate alignment question... ...so after having a read through your reference, the money-shot: Intentional hallucinations primarily happen when o1-preview is asked to provide references to articles, websites, books, or similar sources that it cannot easily verify without access to internet search, causing o1-preview to make up plausible examples instead.
- PoignardAzur 2y agoThe HN discussions for these kinds of articles is so annoying. A third of the discussion follows a pattern of people re-asserting their belief that LLMs can't possibly have knowledge and almost bragging about how they'll ignore any evidence pointing in another direction. They'll ignore it because computers can't possibly understand things in a "real" way and anyone seriously considering the opposite must be deluded about what intelligence is, and they know better. These discussions are fundamentally sterile. They're not about considering ideas or examining evidence, they're about enforcing orthodoxy. Or rather, complaining very loudly that most people don't tightly adhere to their preferred orthodoxy.
- QuantumGood 2y agoWe're in the "the orthodoxies of conventional wisdom have been established, and only a small percentage of people think beyond that" stage?