5 ms·
There’s a few subtle misconceptions being spread here: 1) Hallucination rate is not inversely proportional to number of samples, unless you assume statistical
by refibrillator 2y ago
There’s a few subtle misconceptions being spread here:
1) Hallucination rate is not inversely proportional to number of samples, unless you assume statistical independence. As you’re sampling from the same generative process each time, any inherent bias of the LLM could affect every sample (eg see golden gate Claude). Naively calculating hallucination rate as P^N is going to be a massive underestimate of the true error rate for many tasks requiring factual accuracy.
2) You’re right that output tokens are generated autoregressively, but you are thinking like a human. Transformer attention layers are permutation invariant. The ordering of output (eg decision first then justification later) is inconsequential, either can be derived from input context and hidden state where there is no causal masking of attention.
- bunderbunder 2y agoJustification before decision still works out better in practice, though, because of chain of thought [1]. You'll tend to get more accurate and better-justified decisions. With decision before justification, you tend to have a greater risk of the output being a wrong decision followed by convincing BS justifying it. (edit: Another way you could think of it is, LLMs still can't violate causality. Attention heads' ability to look in both directions with respect to a particular token's position in the sequence does not enable them to see into the future and observe tokens that don't exist yet.) 1: https://arxiv.org/abs/2201.11903 https://arxiv.org/abs/2201.11903
- wtarreau 2y agoI totally agree, that's what I had to do with my patchbot that evaluates haproxy patches to be backported ( https://github.com/haproxy/haproxy/tree/master/dev/patchbot/ https://github.com/haproxy/haproxy/tree/master/dev/patchbot/ ). Originally it would just provide a verdict and justify it and it worked extremely poorly, often with a justification that directly contradicted the verdict. I swapped that by asking the analysis and the final verdict and now the success rate is totally amazing (particularly with mistral that remains unbeatable at this task by obeying extremely well to instructions).
- BeefySwain 2y agoYou find Mistral to be the best "open"/local model? Or you find it to be the best model period?
- deleted 2y ago[deleted]
- solidasparagus 2y agoYour second point I either don’t correctly understand or seems to fly in the face of a lot of proven techniques. Chain-of-thought, react, decision transformer all showcase that order of output of an LLM matter because the tokens output by the LLM before the “answer” can nudge the model to sample from a higher quality part of the distribution for the remainder of its output