4 ms·
Everyone in the comments seems to be arguing over the semantics of the words and anthropomorphization of LLMs. Putting that aside, there is a real problem with
by program_whiz 2y ago
Everyone in the comments seems to be arguing over the semantics of the words and anthropomorphization of LLMs. Putting that aside, there is a real problem with this approach that lies at the mathematical level.
For any given input text, there is a corresponding output text distribution (e.g. the probabilities of all words in a sequence which the model draws samples from).
The approach of drawing several samples and evaluating the entropy and/or disagreement between those draws is that it relies on already knowing the properties of the output distribution. It may be legitimate that one distribution is much more uniformly random than another, which has high certainty. Its not clear to me that they have demonstrated the underlying assumption.
Take for example celebrity info, "What is Tom Cruise known for?". The phrases "movie star", "katie holmes", "topgun", and "scientology" are all quite different in terms of their location in the word vector space, and would result in low semantic similarity, but are all accurate outputs.
On the other hand, "What is Taylor Swift known for?" the answers "standup comedy", "comedian", and "comedy actress" are semantically similar but represent hallucinations. Without knowing the distribution characteristics (e.g multivariate moments and estimates) we couldn't say for certain these are correct merely by their proximity in vector space.
As some have pointed out in this thread, knowing the correct distribution of word sequences for a given input sequence is the very job the LLM is solving, so there is no way of evaluating the output distribution to determine its correctness.
There are actual statistical models to evaluate the amount of uncertainty in output from ANNs (albeit a bit limited), but they are probably not feasible at the scale of LLMs. Perhaps a layer or two could be used to create a partial estimate of uncertainty (e.g. final 2 layers), but this would be a severe truncation of overall network uncertainty.
Another reason I mention this is most hallucinations I encounter are very plausible and often close to the right thing (swapping a variable name, confabulating a config key), which appear very convincing and "in sample", but are actually incorrect.
- byteknight 2y agoYou seem to have explained in much more technical terms than what my "Computer-engineering-without-maths" brain tells me. To me this sounds very similar to lowering temperature. It doesn't sound like it pulls better from grounded-truth but rather more probabilistic in the vector space. Does this jive?
- program_whiz 2y agoPerhaps another way to phrase this is "sampling and evaluating the similarity of samples can determine the dispersion of a distribution, but not its correctness." I can sample a gaussian and tell you how sparse the samples are (standard deviation) but this in no way tells me whether the distribution is accurate (it is possible to have a highly accurate distribution of a high-entropy variable). On the other hand, its possible to have a tight distribution with low standard deviation that is simply inaccurate, but I can't know that simply by sampling from it (unless I already know apriori what the output should look like).
- svnt 2y ago> On the other hand, "What is Taylor Swift known for?" the answers "standup comedy", "comedian", and "comedy actress" are semantically similar but represent hallucinations. Without knowing the distribution characteristics (e.g multivariate moments and estimates) we couldn't say for certain these are correct merely by their proximity in vector space. It depends on the fact that a high uncertainty answer by definition is less probable. That means if you ask multiple times you will not get the same unlikely answer, such as that Taylor swift is a comedian, you will instead get several semantically different answers. Maybe you’re saying the same thing, but if so I’m missing the problem. If your training data tells you that Taylor Swift is known as a comedian, then hallucinations are not your problem.
- sigmoid10 2y agoThis. For a model to consistently output that Taylor Swift is a comedian or something similarly wrong at reasonable temperature settings, there must be a problem in the training data. That doesn't mean that "Taylor Swift is a comedian" needs to be in the training data, it it can simply mean that "Taylor Swift" doesn't appear at all. Then "singer" and "comedian" (and tons of other options) will likely appear at similar probabilities during generation. Blaming human semantics for LLMs is generally a bad idea, since we only use human semantics to qualitatively explain how the models abstract ideas. In practice you simply don't know how the model relates words.
- program_whiz 2y agoIt was just a contrived example to illustrate low variance in the response distribution doesn't necessarily indicate accuracy. Just indicating that "hallucination" is a different axis from "generates different responses" though they might not be totally orthogonal. A better example might be that the model overtrained on AWS cloud formation API 2 and when v3 comes out produces low entropy answers that are wrong for v3 but right for v2 (due to training bias), but the answers are low variance (e.g. "bucket" instead of the new "bucket_name" key). Another example based on a quick test I did on GPT4: In a single phrase, what is Paris? Paris is the city of Light. Paris is the capital of France. Paris is the romantic capital of the world renowned for its art, fashion, and culture.
- kick_in_the_dor 2y agoI think you make a good point, but my guess is that e.g. your Taylor Swift example, a well-grounded model would have a low likelihood of outputting multiple consecutive answers about her being a comedian, which isn't grounded in the training data. For your Tom Cruise example, since all those phrases are true and grounded in the training data, the technique may fire off a false positive "hallucination decision". However, the example they give in the paper seems to be for "single-answer" questions, e.g., "What is the receptor that this very specific medication acts on?", or "Where is the Eiffel Tower located?", in which case I think this approach could be helpful. So perhaps this technique is best-suited for those single-answer applications.
- dwighttk 2y agoWhat’s the single-answer for where the Eiffel Tower is located?
- dontlikeyoueith 2y agoThe Milky Way Galaxy.
- bhaney 2y ago48.8582° N, 2.2945° E
- PeterCorless 2y ago> On the other hand, "What is Taylor Swift known for?" the answers "standup comedy", "comedian", and "comedy actress" are semantically similar but represent hallucinations. Taylor Swift has appeared multiple times on SNL, both as a host and as a surprise guest, beyond being a musical performer[0]. Generally, your point is correct, but she has appeared on the most famous American television show for sketch comedy, making jokes. One can argue whether she was funny or not in her appearances, but she has performed as a comedian, per se. Though she hasn't done a full-on comedy show, she has appeared in comedies in many credits (often as herself).[1] For example she appeared as "Elaine" in a single episode of The New Girl [2x25, "Elaine's Big Day," 2013][2]. She also appeared as Liz Meekins in "Amsterdam" [2022], a black comedy, during which her character is murdered.[3] It'd be interesting if there's such a thing as a negatory hallucination, or, more correctly, an amnesia — the erasure of truth that the AI (for whatever reason) would ignore or discount. [0] https://www.billboard.com/lists/taylor-swift-saturday-night-live-appearances-timeline/ https://www.billboard.com/lists/taylor-swift-saturday-night-... [1] https://www.imdb.com/name/nm2357847/ https://www.imdb.com/name/nm2357847/ [2] https://newgirl.fandom.com/wiki/Elaine https://newgirl.fandom.com/wiki/Elaine [3] https://www.imdb.com/title/tt10304142/?ref_=nm_flmg_t_7_act https://www.imdb.com/title/tt10304142/?ref_=nm_flmg_t_7_act
- gqcwwjtg 2y agoThat doesn’t make it right to say she’s well known for being a comedian.
- leptons 2y agoGarbage in, garbage out. If the "training data" is scraped from online Taylor Swift forums, where her fans are commenting about something funny she did "OMG Taytay is so funny!" "She's hilarious" "She made me laugh so hard" - then the LLM is going to sometimes report that Taylor Swift is a comedian. It's really as simple as that. It's not "hallucinating", it's probability. And it gets worse with AIs being trained on data from reddit and other unreliable sources, where misinformation and disinformation get promoted regularly.
- eutropia 2y agothe method described by this paper does not > draw[ing] several samples and evaluating the entropy and/or disagreement between those draws the method from the paper (as I understand it): - samples multiple answers, (e.g. "music:0.8, musician:0.9, concert:0.7, actress:0.5, superbowl:0.6") - groups them by semantic similarity and gives them an id ([music, musician, concert] -> MUSIC, [actress] -> ACTING, [superbowl] -> SPORTS), note that they just use an integer or something for the id - sums the probability of those grouped answers and normalizes: (MUSIC:2.4, ACTING:0.5, SPORTS:0.6 -> MUSIC:0.686, SPORTS:0.171, ACTING:0.143) They also go to pains in the paper to clearly define what they are trying to prevent, which is confabulations. > We focus on a subset of hallucinations which we call ‘confabulations’ for which LLMs fluently make claims that are both wrong and arbitrary—by which we mean that the answer is sensitive to irrelevant details such as random seed. Common misconceptions will still be strongly represented in the dataset. What this method does is it penalizes semantically isolated answers (answers dissimilar to other possible answers) with mediocre likelihood. Now technically, this paper only compares the effectiveness of "detecting" the confabulation to other methods - it doesn't offer an improved sampling method which utilizes that detection. And of course, if it were used as part of a generation technique it is subject to the extreme penalty of 10xing the number of model generations required. link to the code: https://github.com/jlko/semantic_uncertainty https://github.com/jlko/semantic_uncertainty
- eigenspace 2y agoRight, but the problem pointed out here is that if you compare that to the answers for Tom Cruise, you’d get a bunch of disparate answers that under this method would seem to indicate that it was confabulating, when in reality, Tom Cruise is just known for a lot of different things.
- program_whiz 2y agoI don't think discretizing the results solves the problem, we don't know whether the distribution is accurate without apriori knowledge. See my real GTP4 output about Paris. Are the words "city of light" "center of culture" and "capital of France" confabulations? Without apriori knowledge is it more or less confabulatory than "city of roses", "site of religious significance", "capital of Korea"? If it simply output "Capital of Rome" 3 times, would that indicate its probably not a confabulation? You can discretize the concepts but that only serves to reduce the granularity of comparisons, and does solve the underlying problem I originally described.
- bubblyworld 2y agoThey're not using vector embeddings for determining similarity - they use finetuned NLI models that take the context into account to determine semantic equivalence. So it doesn't depend on knowing properties of the output distribution up front at all. All you need to be able to do is draw a representative sample (up to your preferred error bounds).