4 ms·
Some of the comments reminded me of LeCun's claim regarding the error distribution of an LLM output conditional on content length. Namely, if "e" is the probabi
by usgroup 3y ago
Some of the comments reminded me of LeCun's claim regarding the error distribution of an LLM output conditional on content length. Namely, if "e" is the probability of an error, the probability of a sequence of length "n" being error free is p = (1-e)^n. That is to say there is exponentially less chance that an LLM sequence is "within the distribution of correct answers" as token length increases.
This is a consequence of the "auto-regressive" model and its lack of in-built self-correction, and it is a limiting factor in actual applications.
LeCun's tweet:
https://twitter.com/ylecun/status/1640122342570336267 https://twitter.com/ylecun/status/1640122342570336267