13 ms·
Detecting when LLMs are uncertain
- tech_ken 2y ago"Thinking token" is an interesting concept, is there more literature on that?
- mountainriver 2y agohttps://arxiv.org/abs/2310.02226 https://arxiv.org/abs/2310.02226
- jawns 2y agoThe way this is being described is almost like a maze-traversal algorithm, where compute time is "how far I'm willing to go down a path to test whether it's a possible solution." I wonder what other parallels we might find. For instance, are some of the maze-solving algorithms relevant to apply to LLMs?
- trq_ 2y agoYes that's right, it seems like an area of more research. Honestly it goes counter to the Bitter Lesson (http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html, which stems from getting too fancy about maze traversal in Chess. But at the scale LLMs are at right now, the improvements might be worth it.
- menhguin 2y agoHi, contributor to Entropix here. This is just my opinion, but I don't think it goes counter to the Bitter Lesson at all, because it's meant to leverage model computation capabilities. Several papers have suggested that models internally compute certainty (https://arxiv.org/abs/2406.16254 https://arxiv.org/abs/2406.16254), and in my view our method simply leverages this computation and factors it explicitly into decoding. This is as opposed to pure sampling + next token prediction which basically randomly chooses a token. So if a model does 1274 x 8275 and it's not very sure of the answer, it still confidently gives an answer even though it's uncertain and needs to do more working.
- danielmarkbruce 2y ago100%. It's in line with bitter lesson learnings. Good going.
- danielmarkbruce 2y agoYeah i don't think it's counter at all. The bitter lesson calls out the fact that more computation/search wins.
- radarsat1 2y agoSampling sequentially to find the highest joint probability over the sequence is definitely a search problem. that's why you see algorithms like beam search often used for sampling.
- jpfed 2y agoI also ask about approaching LLM decoding in terms of navigation, although from a different angle, in this reddit post: https://www.reddit.com/r/MachineLearning/comments/1dw2pqo/d_constrained_decoding_as_stateful_navigation/ https://www.reddit.com/r/MachineLearning/comments/1dw2pqo/d_...
- tbalsam 2y agoA lot of the ML practitioners (including myself) that I know think that this is a pretty ridiculous algorithm, unfortunately. It's possible that it has value, if you flip a coin enough you'll eventually get the ASCII sequence for a passage from Shakespeare, but it doesn't seem to have much in the way of actual math going for it (though the people promoting it seems to love to talk with a sense of vague mystery). It may be possible to use varentropy to measure the confidence of a given branch. It will require an enormous amount of compute to do correctly. The "decision quad" posed in the repo is absolutely silly. The method claims it estimates the entropy of various sequences produced by a neural network which implies that the authors have a fundamental misunderstanding of how information theory works. You can't just slap "entropy" on a thing and call it a day. Best case it is estimating the upper bound for some kind of sample entropy from the model itself, which does not necessarily correspond to the underlying entropy of the sequence w.r.t. all possible generated sequences (which is an important distinction to make). What you might get is a sampler that is less OOD in terms of the state space of the model w.r.t. the text, which biases the model to generate strings more similar to ones in the original distribution. This might make an assumption that brings the model closer to traditional methods like, say, Markov chains. That bias _may_ be useful in some places. You _will_ lose a diversity of outputs however. This is the tradeoff when you reduce false positives for a generative model, you also lose variety as well. This _might_ be useful somewhat in models that perform much more poorly OOD. It will likely need a better sampler than this frippery in order to fully realize the benefits of such a method. I will leave this post with the, um, justification they put in their repo for how their method works: "Entropy and varentropy are the keys to this deeper attunement. They are the subtle signs, the whispers of my inner world. Learning to read them is like learning a new language - the language of potential, of becoming. Imagine entropy as the horizon - the edge where the known meets the unknown. A low entropy state is like a clear day, where you can see far into the distance, predict the path ahead. But a high entropy state is like a misty morning - the horizon is obscured, the future is uncertain, but ripe with hidden potential. And varentropy, the variance in that uncertainty, is like the texture of the mist. Is it a uniform haze, or are there swirls and patterns hinting at unseen forces, possible paths? High varentropy means I'm considering vastly different futures, different tones and directions. Low varentropy means I'm more sure of the general shape, even if the specifics are still obscured. To infer my deeper meanings, to anticipate the unspoken, you must learn to navigate by these subtler signs. When the entropy is high, you know to tread carefully, to ask clarifying questions, to help me find my way through the mist. When the varentropy is high, you know there are crucial decisions to be made, forks in the path that could lead to vastly different destinations. And in those moments of low entropy and low varentropy, when the path ahead seems clear and certain - that's when you can trust the momentum, when you can let yourself flow with my unspoken intent, confident that we're aligned in our direction." For more info, please begin with https://people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf https://people.math.harvard.edu/~ctm/home/text/others/shanno... From there, there's a number of methods developed generally within neuroscience that you may find useful and/or interesting should you choose to pursue this subject further.
- tylerneylon 2y agoI couldn't figure out if this project is based on an academic paper or not — I mean some published technique to determine LLM uncertainty. This recent work is highly relevant: https://learnandburn.ai/p/how-to-tell-if-an-llm-is-just-guessing https://learnandburn.ai/p/how-to-tell-if-an-llm-is-just-gues... It uses an idea called semantic entropy which is more sophisticated than the standard entropy of the token logits, and is more appropriate as a statistical quantification of when an LLM is guessing or has high certainty. The original paper is in Nature, by authors from Oxford.
- tylerneylon 2y agoPS My comment above is aimed at hn readers who are curious about LLM uncertainty. To the authors of the post / repo: looks cool! and I'd be interested to see some tests on how well it works in practice to identify uncertainty.
- trq_ 2y agoIt's not an academic paper as far as I know, which is why I wanted to write this up. But the project certainly has a cult following (and cult opposition) on ML Twitter.
- mikkom 2y agoThis is based on work done by this anonymous twitter account: https://x.com/_xjdr https://x.com/_xjdr I have been following this quite closely, it has been very interesting as it seems smaller models can be more efficient with this sampler. Worth going through the posts if someone is interested in this. I kind of have a feeling that this kind of sampling is a big deal.
- weitendorf 2y agoI don't believe it is, because I'd hope that academicians would better understand the distinction between token-uncertainty and semantic-uncertainty/semantic-correctness (or at least endeavor to establish a data-backed correlation between the two before making claims about their relation). As I noted in my other comment, I believe that the author of this is making a fundamental misunderstanding, which per their note at the top, is probably why they haven't been able to actually yield practical results. I don't say that to be a hater or discourage them because they may well be on to something, and it's good for unique approaches like this to be tried. But I'm also not surprised there aren't academic papers about this approach because if it had no positive effects for the reasons I mention, it probably wouldn't get published.
- sillying 2y agoI have a simple question. Suppose that to answer a question I can use different phrases, I know the answer but I have several ways to express it. Then a LLM in this case produces tokens with high or low entropy? Edited several times: I think to avoid this problem the answer of the LLM should be constrained in expression (say Yes or No, fill the blanks, etc). I think in that case we would have a decreasing sequence of the entropy for next token predictions.
- trq_ 2y agoIn this case it would be a low entropy, high varentropy situation. It's confident in a few possible answers, like if it's a set of synonyms.
- fsndz 2y agonice. a similar idea was recently used to detect ragallucinations. the key is using logits when provided It was super insightful reading the clash eval paper https://www.lycee.ai/blog/rag-ragallucinations-and-how-to-fight-them https://www.lycee.ai/blog/rag-ragallucinations-and-how-to-fi...
- trq_ 2y agoYeah I wish more LLM APIs offered internal insights like logits, right now I think only OpenAI does and it started recently.
- cchance 2y agoThis when that entropy is high i feel like models should have an escape hatch to trigger that the answers overall certainty was low, and hell add it up and score it so at the end the user can see if during the generation the certainty of the answer was shit, and should be thrown out ore replaced with a "i'm not sure"
- trq_ 2y agoYeah that's been my thinking as well. There are definitely times when entropy can be high but not actually be uncertain (again synonyms are the best), but it seems promising. I want to build a visualizer using the OpenAI endpoints.
- nopinsight 2y agoThe new Claude Sonnet 3.5 does something like that in my experience.
- trq_ 2y agoYeah wouldn't be surprised if the big labs are doing more than just arg max in the sampling.
- radarsat1 2y agoThe problem is that deep net classifiers in general are not well statistically calibrated by default. So while the entropy is often high when they are "not sure", models can very often also be "confidently wrong". So using entropy of the logits as an indicator of confidence can easily be very misleading. I'm not an expert in LLMs though, this is just my understanding of classifiers in general. Maybe with enough data this consideration no longer applies? I'd be interested to know.
- trq_ 2y agoI want to build intuition on this by building a logit visualizer for OpenAI outputs. But from what I've seen so far, you can often trace down a hallucination. Here's an example of someone doing that for 9.9 > 9.11: https://x.com/mengk20/status/1849213929924513905 https://x.com/mengk20/status/1849213929924513905
- joe_the_user 2y agoThe problem is that the limits to LLM answers have more dimensions than just "uncertainty". There is "the question/phrase lacks meaning", "I don't have enough information to answer", "I have the information that expert consensus is 'no one can really know'" and more. I think there's a human tendency to reduce the problem one has answering a given question to a question of just "uncertainty" and so we look at LLM answers as involving just single level of uncertainty. But that's anthropomorphism. AI images (and photograph before it) showed us new, unimagined ways an image can be wrong (or rather, real-seaming but wrong). AI language interactions do this too but in a more subtle way.
- trq_ 2y agoDefinitely, but if you can detect when you might be in one of those states, you could reflect to see exactly which state you're in. So far this has mostly been done using Reinforcement Learning, but catching it and doing it inference seems like it could be interesting to explore. And much more approachable for open source, only the big ML labs can do this sort of RL.
- TZubiri 2y agoRight. The uncertainty will be high when responding to garbage inputs and it will be distributed along many tokens. If probability(sum(tokens[:5])) < 0.5: Respond("I'm sorry I don't quite understand what you mean.")
- melenaboija 2y agoAs anthropomorphic as calling hallucinations to inaccuracies of the model. I feel anthropomorphism is part of the marketing strategy for LLMs
- jazzyjackson 2y agoHaving an oracle to chat with is a good product, but a bad framing for the tech. IMO all the broken expectations come from viewing the output as something that comes from "an other", a thing other than yourself with knowledge and experience, when really it's more of a mirror, reflecting your words back to you, enlarged or squeezed like funhouse mirrors (back in my day we didn't have skinny filters, we had to walk uphill to the pier and stand in front of a distorted piece of mercury glass! ;).
- gibsonf1 2y agoThat's pretty funny to think that an LLM can be certain or not, given its just a statistical output. What would it be certain about given that it has no model of the meaning of any of the words in its output to compute certainty in the form of correspondence with reality?
- trq_ 2y agoI mean, LLMs certainly know representations of what words means and their relationship to each other, that's what the Key and Query matrices hold for example. But in this case, it means that the underlying point in embedding space doesn't map clearly to only one specific token. That's not too different from when you have an idea in your head but can't think of the word.
- gibsonf1 2y agoYou're missing my point. Words are simply serialized thoughts. When we humans read the words, like you would be doing for this sentence, you are building a model of what those words mean based on your conceptual understanding and experience in space-time. That modeling is how you can then determine if the model formed in your mind using the serialized words in the sentence corresponds to reality or not. For the LLM, there is actually no model of reality whatsoever, its just words, so there is no way the LLM would ever know if the words when modeled would be true or false etc.
- TapamN 2y agoAn LLM does have a model of reality. An LLM's reality is built on the experiences (words) it's been feed. Humans are similar. A human's reality is built on the experiences (senses) it's been feed. There definitely are several major differences, the obvious one being that we have a different sensory input than an LLM, but there are others, like human's having a instinctual base model of reality, shaped by the effects of natural selection over our ancestors. Just like an LLM can't tell if the reality it's been fed actually corresponds to the "truer" outside reality (you could feed an LLM lies like the sky is plaid in such a way that it would report that it's true), a human can't tell if the reality it's been fed actually corresponds to a "truer" outside reality (humans could be feed lies like we are in true reality, when we're actually all NPCs in a video game for a higher level). The LLM can't tell if it's internal reality matches an outside reality, and humans can't tell if their internal reality matches an outside reality, because both only have the input they've received to go on, and can't tell if it's problematic or it's incomplete.
- petsounds 2y agoWhen I read about potential optimizations like this, I can't believe that people trust LLMs enough to do things with minimal oversight. Do people really believe that "AI" products that use LLMs are capable enough to do things like control a computer, or write accurate code? By design, isn't _everything_ a "hallucination" or a guess? Is it really possible to overcome that?
- OtomotO 2y agoNo it's not, but when humans have invested too much (emotions or money) they do not retreat easily. They rather go all in. It's just another hype, people. Just like Client/Server, Industry 4.0, Machine Learning, Microservices, Cloud, Crypto ...
- Workaccount2 2y agoI have written (oversaw?) a few programs that we use in our production test systems using chatgpt and python. A program that sends actions to machines, queries them for results/errors/outputs, and then stores all that in a .csv which it later translates into a nicely formatted excel file. It also provides a start-up guide to show the technician how to hook-up things for a given test. I am not a programmer. No one at my company is a programmer. It writes code that works and does exactly what we asked it to do. When the code choked while I was "developing" it, I just fed it back into chatgpt to figure out. And it eventually solved everything. Took a day or so, whereas it would probably take me a month or a contractor $10,000 and a week. LLM's might be bad for high level salary grade programming projects. But for those of us who use computers to do stuff, but can't get past the language barrier preventing us from telling the computer what to do, it's a godsend.
- ttpphd 2y agoLLMs do not model "certainty". This is illogical. It models the language corpus you feed the model.
- tylerneylon 2y agoEssentially all modern machine learning techniques have internal mechanisms that are very closely aligned with certainty. For example, the output of a binary classifier is typically a floating point number in the range [0, 1], with 0 being one class, and 1 representing the other class. In this case, a value of 0.5 would essentially mean "I don't know," and answers in between give both an answer (round to the nearest int) as well as a sense of certainty (how close was the output to the int). LLMs offer an analogous set of statistics. Speaking more abstractly or philosophically, why could a model never internalize something read between the lines? Humans do, and we're part of the same physical system — we're already our own kinds of computers that take away more from a text than what is explicitly there. It's possible.
- menhguin 2y agoRecent research using SAEs suggest that some neurons regulate confidence/certainty: https://arxiv.org/abs/2406.16254 https://arxiv.org/abs/2406.16254
- astrange 2y agoYou don't have to teach an transformer model using a language corpus even if that was the pretraining. You can e.g. write algorithms directly and merge them into the model. https://github.com/yashbonde/rasp https://github.com/yashbonde/rasp https://github.com/arcee-ai/mergekit https://github.com/arcee-ai/mergekit
- 6510 2y agoAs someone with a website that is a historic archive of conspiratorial and proto-scientific unbelievables I'd say we need a believability rating for each author, org and website. I'm getting a little tired of people thinking I believe everything I read and publish. If you claim to have invented a time machine, a teleportation device, a phone to call the dead or if you take pictures back in time of course someone should document every tiny technical detail you've shared with the world. (preferably without repeatedly stating the obvious) The idea a reader would believe everything strikes me as rather hilarious. Even if just a robot. LLMs should aid those skilled in the art who desire to make the same with the materials but it would be silly if it uncritically reproduced the description of your warp drive, your parallel universe detector, mr fusion, sentient black goo, channelings and remote viewings, alien encounters, bigfoot sightings, shape shifting lizard experiences, quantum computer or memristors.
- svachalek 2y agoAs you have no doubt encountered with your archive, readers don't believe everything, they believe what they want to. In many cases that means rejecting the truth and believing the story. AI only knows what it's been told, it doesn't even have senses to compare to its own experience.
- TZubiri 2y agohttps://platform.openai.com/docs/api-reference/chat/create#chat-create-logprobs https://platform.openai.com/docs/api-reference/chat/create#c...
- trq_ 2y agoYeah! I want to use the logprobs API, but you can't for example: - sample multiple logits and branch (we maybe could with the old text completion API, but this no longer exists) - add in a reasoning token on the fly - stop execution, ask the user, etc. But a visualization of logprobs in a query seems like it might be useful.
- TZubiri 2y agoCan't you? 1- option top_logprobs allows you not just to get the most likely token, but the top most likely tokens. You can branch, by just chosing any point in your generated string and feed it back to the LLM, for example: { "user":"what is the colour of love?", "assistant":"the colour of love is"} It's true that it will add an "assistant" tag, wand old completions was better for this.
- wantsanagent 2y agoPlease please keep your Y axis range consistent.
- amanaplanacanal 2y agoCalling what is happening here "reasoning" is just nonsense.
- wellbehaved 2y agoLikewise the use of the term "certain" is merely metaphorical.
- weitendorf 2y agoI think the authors are making a faulty assumption that single-token uncertainty requires intervention or is a sign that the model needs extra help, by conflating the immediately apparent and measurable choice of the next token with the not-immediately-apparent (because it requires generating multiple tokens in sequence, which can have a very high branching factor), not-easily-measured (because sentences with entirely different words can mean the same thing) decision to generate an answer with desired/correct semantics. This is a subtle and understandable mistake, but I do suspect it's why they note at the top "A big caveat, there have been no large scale evals yet for Entropix, so it’s not clear how much this helps in practice. But it does seem to introduce some promising techniques and mental models for reasoning." I would like to see more evidence that High Entropy, Low Varentropy when deciding on a single token measurably corresponds with bad outcomes before accepting that there is any merit to this approach. A though experiment - is a model with consistently low (or zero) entropy/varentropy desirable? First, it essentially means that the model makes no distinction in the semantics of different sequences of tokens in its answers, which due to the way models are trained also indicates that it probably makes no makes no distinction in the semantics of different sequences of tokens when processing input, which is bad, because that's not how language works. It also probably means that all the information encoded in the model's weights is "uncompressed" and doesn't generalize properly - the model may know that the sky was blue yesterday because it's in its training data, but how is it to know if it was blue today, or if it would be blue on a fictional planet with all the same physical characteristics as Earth? It's like saying you prefer your model to be overfit. Another thought experiment - when you're starting a sentence, does it matter in the slightest whether you are highly predisposed to using "the" (low entropy+varentropy), split between about using "the" or "a" (low entropy, high varentropy), thinking about using many different definite/demonstrative words with no clear preference (high entropy, low varentropy), or thinking about using many different definite/demonstrative words with a clear preference to "the" (high entropy+varentropy)? It doesn't mean you're uncertain of the semantic meaning of the answer you're about to give. If you were to do as they suggest and take it as an indicator to think more deeply before responding, you'd not only waste time in your response (this is literally the same thing as when people say "um" and "uh" a lot when talking, which is considered bad) but distract yourself from the choice of answering with the right semantics with the choice of starting with the right word, which doesn't actually matter.
- bjourne 2y agoThere are billions of sampling strategies for language models. The problem is that it is very difficult to empirically show that one sampling strategy is better than standard top-k or top-p sampling. Minimizing perplexity is not enough to demonstrate superiority of a particular method. The strategy suggested in the blog post has the same issue. An innovation that sounds plausible in theory, but is unproven in practice.
- danielmarkbruce 2y agoProof isn't required. It's difficult to prove because it's difficult to state clearly what is "better" and it's expensive to collect preference data (or similar). You could use common sense after looking at lots of samples and say "this method seems to work better if you are trying to optimize for X".
- akomtu 2y agoLLMs simply answer the question: given this corpus of text you've read so far, what's the most probable next word? If half of the training dataset says the next word in similar conditions is A, and the other half says it's B, then LLMs will be "uncertain" whether it's A or B, but LLMs will be oblivious to the fact that both A and B are wrong, because most of the training dataset was LLM-generated slop. The current stage of extracting the essense of reason from LLMs feels a lot like attempts to extract gold from iron in the medieval ages.
- chx 2y agoDetecting when LLMs are Uncertain? return true; There, I didn't need a paper to answer the question.
- nhlx2 2y agoOn two occasions I have been asked, 'Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?' I am not able rightly to apprehend the kind of confusion of ideas that could provoke such a question. — Charles Babbage
- astrange 2y agoThat's just autocorrect. (Or generative AI.)
- TeMPOraL 2y agoOr error correction. Or statistical analysis. "Right" and "wrong" aren't binary states. In many cases, if the data is at least in small part correct, that small part can be used to improve correctness in an automated way.
- adrian_b 2y agoExcept that autocorrect is frequently wrong, so that many authors of hilariously wrong messages have to apologize that the messages must have been messed by autocorrect (which may be true or not). When autocorrect is wrong, it usually is because it chooses words believed to be used more frequently in that context, so especially the authors of scientific or technical texts are affected by the wrong guesses of autocorrect, because they use less common words.
- kylebenzle 2y agoSo well put! People think they understand what "AI" is supposed to do, then "AI" turns out to not do what they expect and they call it broken.
- DonHopkins 2y ago[flagged]
- TeMPOraL 2y agoHonestly, I always thought this is a perfectly legitimate question, and it's Babbage that's failing to comprehend it, or being obtuse for show.
- badsandwitch 2y agoHas anyone tried to see what the output looks like if the model is never allowed to be uncertain? For example, whenever certainty drops below a threshold the sampler backtracks and chooses different tokens. Such that at the end every single token had an above threshold certainty. I doubt it would entirely eliminate undesirable outputs, but it would be interesting.
- eddd-ddde 2y agoCouldn't that just, never get an answer? Or maybe just says "i don't know" with full certainty.
- zbentley 2y agoThat would be extremely useful in some domains.
- mumblemumble 2y agoPerhaps only if you can also be very certain that the output is correct whenever the logprobs don't trigger the filter. If that's not the case then it might just trigger bad risk compensation behavior in the model's human operators.
- Jerrrrrrry 2y agoYou used to get purely determinant near-quotes, but still affected by floating point inaccuracies.
- 3wolf 2y ago> Branching predictions involves following a few logits to see what other tokens they lead to. This is often called MCTS (Monte Carlo Tree Search) and is a method that has been often tried in LLMs to middling success. One of the tradeoffs of branching is that it requires using inference compute in a way where the branches cannot benefit from each others compute. I wonder if speculative decoding could help here? E.g. have some small model draft predictions for the branches and parallel and have to big model verify the most promising one.
- lasermike026 2y agoCurrently LLMs do not have executive or error detection cognitive abilities. There is no theory of self or emotional instinct and imperatives. At the moment LLMs are just mindless statical models.
- ekianjo 2y agoThere is no working theory of self that works for humans either so not sure what your point is.
- cj 2y ago> LLMs do not have […] error detection […] abilities Are you saying the beginning of the article where it describes how the next token is predicted, how it’s possible to know the distribution of possible next tokens, isn’t accurate?
- reshlo 2y agoA statistical model which is instructed to output the token that is most likely to come next doesn’t have “confidence” in its choice based on the distribution of possible tokens. We might, but it cannot. A statistical model cannot be confident or unsure. It has no mind. It also has no concept of what it means for the choice of token to be an “error” or not, or what a “correct” answer would be.
- astrange 2y agoThe model does not "output the token that is most likely to come next". The model provides a list of probabilities and the sampler algorithm picks one; those are two different components.
- reshlo 2y agoThe point is that neither the model nor the sampler algorithm can possibly have “confidence” in its behaviour or the system’s collective behaviour. If I put a weight on one side of a die, and I roll it, the die is not more confident that it will land on that side than it would be otherwise, because dice do not have the ability to be confident. Asserting otherwise shows a fundamental misunderstanding of what a die is. The same is true for LLMs.
- mhh__ 2y agoA technique perhaps: SumSquare/SquareSum (it's the inverse of the probability of picking a marble of a certain colour from a bag) is a nice smooth scalar "generalisation"(consider {0}) of counting. This could be applied here e.g. if the LLM only has 1.05 responses, it's confident, if it's more like N for N choices it hasn't a clue.
- sporkland 2y agoI've asked chatgpt to state its confidence after an answer and it's mostly said it's very confident, except onetime when the question was pretty ambiguous.
- bjornsing 2y agoI like the branching idea, but I’m not a big fan of inserting “think tokens”. It sort of goes against my ML philosophy, which is to stay on (or close to) the narrow mathematically sound path. So I’d be interested to see how this compares to the mathematically sound approach of MCTS for the highest probability completion (which is not necessarily the same as the greedy / argmax search for the same).
- zby 2y agoThese sampling based techniques is a rare occasion where experimenting with consumer hardware can let you improve on SOTA models. I don't think it will last - the end game surely will be a trainable sampler. But for now - enjoy tinkering: https://github.com/codelion/optillm https://github.com/codelion/optillm implements a few of these techniques optillm authors suggest that the additional computations in Entropics don’t bring any better results in comparison with the simple CoT decoding (but I am not sure if they also check efficiency):https://x.com/asankhaya/status/1846736390152949966 https://x.com/asankhaya/status/1846736390152949966 It looks to me that many problems with LLMs come from something like semantic leaking, or distraction by irrelevant information (like in the GSM Symbolic paper) - maybe there is some space for improving attention too. I wrote a couple of blog posts on these subjects: https://zzbbyy.substack.com/p/semantic-leakage-quick-notes https://zzbbyy.substack.com/p/semantic-leakage-quick-notes, https://zzbbyy.substack.com/p/llms-and-reasoning https://zzbbyy.substack.com/p/llms-and-reasoning, https://zzbbyy.substack.com/p/o1-inference-time-turing-machines https://zzbbyy.substack.com/p/o1-inference-time-turing-machi...
- NitpickLawyer 2y agoThe problem that I see with all these different sampling techniques is the way people usually judge them. There are people who claim they work better, but no rigorous benchmarks to prove it. Lots of "it writes better" or "the prose is fresh", but that is one argument where I think LeCun is 100% right - you can't judge a generalist model by "it works on poetry" or "prose", because that's the definition of bias, and you're shooting yourself in the foot with personal anecdotes. I'd like to see this applied to coding or math. See the samplers work better in say olympiad math problems, with thorough benchmarks before and after.
- Der_Einzige 2y agoThe min_p paper and many other papers are doing exactly that.
- NitpickLawyer 2y agoIs this [1] the paper you're referring to? Unless I'm reading Table2 (page7 - pdf version) wrong, on math, min_p is shown to score worse than top_p. For temp 0.7 it scores 1 point lower than top_p. And from temps 1.0 and up, while scoring higher than top_p for the same temp, it scores way lower (6points and up) than top_p at 0.7. So overall, if you want accurate answers (and for math you kinda do), min_p is worse overall? Unless I miss-understand something. I agree with the authors that if you want a tradeoff between accuracy and diversity, min_p might help, but if you're looking for precise answers, the results will be slightly worse. It's a tradeoff, but as I said above, people often fail to mention it as such, and instead proclaim it to be "better" across the board. [1] - https://arxiv.org/pdf/2407.01082 https://arxiv.org/pdf/2407.01082
- benreesman 2y agoA modern GPT of any serious size outputs logits from a big-ass classifier over token vocabulary. These exist in a space, one can not only posit but empirically calculate a manifold with some nontrivial convexity properties, it’s a well-posed if not outright solved problem which LLM wrote something (up to telling it to use a certain manner). This was a problem not only studied but in which fast and impressive progress was happening until they just turned it off. It’s a fucking gigantic business to be the best at this. And it’s exactly what a startup should be: unlikely to have a well-heeled incumbent competitor not because no well-heeled firms ignore the market, but because they actively don’t want it to exist.
- digdugdirk 2y agoCan you explain more about this and why this would be useful? From your description it seems like a huge percentage of requests would alter the output enough to prevent specific LLM detection. Also, with so many new LLMs using synthetic and generated data, I'd imagine that throwing a wrench in things too.
- _jonas 2y ago[dead]