3 ms·
It is not semantics. For decades, logic and CS researchers have known what reasoning is. LLM folks suddenly can’t claim an approximation of that is what consti
by thesmtsolver2 2mo ago
It is not semantics. For decades, logic and CS researchers have known what reasoning is.
LLM folks suddenly can’t claim an approximation of that is what constitutes full scale reasoning just because they can achieve only an approximation.
Imagine a calculator program that computes billions of two number multiplications accurately by looking up prior examples but fails on simple multiplications often as it doesn’t have it in its training dataset.
We won’t say the program actually multiplies numbers.
- solid_fuel 2mo ago> Imagine a calculator program that computes billions of two number multiplications accurately by looking up prior examples but fails on simple multiplications often as it doesn’t have it in its training dataset. > We won’t say the program actually multiplies numbers. That's a good analogy. To extend it further, in cases where the calculator can't handle a question - e.g. numbers too large - a properly designed calculator returns an error instead of a randomly hallucinated answer. We haven't even achieved that level of safeguard around token predictors yet.
- urbsgpw 2mo agoBut isn't it the case that we can't reach this safeguard with the current architecture? I remember Karpathy making an interesting point 2 years (cca) back, that I would summarize somehow like this: the mechanics behind every LLM answer are the same, what you then call hallucination is more or less a consequence of whether or not the answer was factually correct/useful. Which would mean, as is so often the case, that the "killer feature" of the LLMs is also its biggest weakness and the two can't be disentangled. Now, we are inventive creatures and we might come up with a remedy for these issues, but what you basically see so far is more guardrails, the use of harnesses and building a whole bunch of infrastructure around the LLMs to get useful work out of them. Which, btw is not a critique, I do it as well and it's a fun engineering challenge.
- simianwords 2mo agoThis is a strawman. LLMs have the ability to identify when they can’t solve a question just like humans.
- solid_fuel 2mo agoA strawman? Don’t be ridiculous. Hallucination remains an unsolved problem, and LLMs do not have any meaningful ability to recognize when a conversation has steered outside the training data. If you can somehow change that, there is a Turing award waiting.
- gr_norm 2mo agoAgree, I've raised this point often. And certainly what remains is still useful, once you accept it! But under no circumstances can we allow scientific achievements to be falsely claimed in service of justifying huge capital investments. Attempting an end-run around the truth, here by redefining words to mean things they don't, always slows down real progress.
- quietbritishjim 2mo ago> Imagine a calculator program that computes billions of two number multiplications accurately by looking up prior examples This is a poor analogy because: * Multiplying numbers has a single objective answer. Whether a code is good (sometimes even just whether it's correct) can be quite subjective. * LLMs certainly do some level of composition between the data sources they were trained on i.e. they are more than just lookup tables. * We have calculators that actually do multiply large numbers accurately. We don't have anything that automatically writes code that is definitely correct and "good". To address only the last point: imagine you had a device that would quickly factor large numbers used in modern criticality, but occasionally got it wrong. You could waste a lot time debating whether it's a "calculator", but it's still certainly useful to have one.
- thesmtsolver2 2mo agoI think you are just agreeing with me and restating the inputs to my argument. I didn’t say LLMs are capable of zero reasoning. An approximation is just an approximation no matter how good. Will you bet your wealth or critical safety systems on: 1. Accuracy of standard computer arithmetic: Yes 2. Accuracy of Lean or automated theorem provers: Yes (you already do) 3. Accuracy of LLM reasoning: No https://en.wikipedia.org/wiki/Handbook_of_Automated_Reasoning https://en.wikipedia.org/wiki/Handbook_of_Automated_Reasonin...
- fragmede 2mo ago> For decades, logic and CS researchers have known what reasoning is. Oh good. Can you share that definition with the rest of us then? We're out here fumbling around trying to define what reasoning is without the benefit of their definition.
- bluejay2387 2mo ago"For decades, logic and CS researchers have known what reasoning is." ... this is a fairly significant overstatement. There is not complete agreement on this term and our understanding continues to evolve. The models don't have to think like humans to think. Saying that LLM's only offer an 'approximation' of reasoning is also an overstatement as it is not a resolved topic. But to the original point, its not exactly just semantics if thought traces are not doing the job that they were originally thought to do. There is value in knowing how these things actually work. If chain of thought is just grounding the latent space and not directly contributing to the process of generating a response it has implications on how we test and verify the reliability of models if nothing else... doesn't mean they aren't useful but it definitely impacts many of the tools we could have to evaluate their performance.
- kevinwang 2mo ago> It is not semantics. For decades, logic and CS researchers have known what reasoning is. Curious what this is!
- paulddraper 2mo agoMe too. What does a reasoning program look like and why is matrix multiplication not that?
- thesmtsolver2 2mo agoThis is a great resource https://en.wikipedia.org/wiki/Handbook_of_Automated_Reasoning https://en.wikipedia.org/wiki/Handbook_of_Automated_Reasonin...
- dumah 2mo agoLots of people asking here for this clear definition. If you understand this well, please lay it out here in a straightforward manner.