4 ms·
> Imagine a calculator program that computes billions of two number multiplications accurately by looking up prior examples but fails on simple multiplications
by solid_fuel 2mo ago
> Imagine a calculator program that computes billions of two number multiplications accurately by looking up prior examples but fails on simple multiplications often as it doesn’t have it in its training dataset.
> We won’t say the program actually multiplies numbers.
That's a good analogy. To extend it further, in cases where the calculator can't handle a question - e.g. numbers too large - a properly designed calculator returns an error instead of a randomly hallucinated answer. We haven't even achieved that level of safeguard around token predictors yet.
- urbsgpw 2mo agoBut isn't it the case that we can't reach this safeguard with the current architecture? I remember Karpathy making an interesting point 2 years (cca) back, that I would summarize somehow like this: the mechanics behind every LLM answer are the same, what you then call hallucination is more or less a consequence of whether or not the answer was factually correct/useful. Which would mean, as is so often the case, that the "killer feature" of the LLMs is also its biggest weakness and the two can't be disentangled. Now, we are inventive creatures and we might come up with a remedy for these issues, but what you basically see so far is more guardrails, the use of harnesses and building a whole bunch of infrastructure around the LLMs to get useful work out of them. Which, btw is not a critique, I do it as well and it's a fun engineering challenge.
- simianwords 2mo agoThis is a strawman. LLMs have the ability to identify when they can’t solve a question just like humans.
- solid_fuel 2mo agoA strawman? Don’t be ridiculous. Hallucination remains an unsolved problem, and LLMs do not have any meaningful ability to recognize when a conversation has steered outside the training data. If you can somehow change that, there is a Turing award waiting.