4 ms·
LLM gets things right, when it does, due to the sheer massive information ingested during training, it can use probabilities to extract a right answer from deep
by ojosilva 2y ago
LLM gets things right, when it does, due to the sheer massive information ingested during training, it can use probabilities to extract a right answer from deep in the model.
Humans on the other hand have developed a more elaborate scheme to process, or reason, data without having to read through 1 billion math problems and stack overflow answers. We listen to some explanations, a YT video, a few exercises and we're ready to go.
The fact that we may get similar grades (at ie high school math) is just a spot coincidence of where both "species" (AI x Human) are right now at succeeding. But if we look closer at failure, we'll see that we fail very differently. AI failure right now looks, to us humans, very nonsensical.
- pishpash 2y agoNah, human failures look equally nonsensical. You're just more attuned to use their body language or peer judgement to augment your reception. Really psychotic humans can bypass this check.
- ben_w 2y agoWhile I'd agree human failures are different from AI failures, human failures are necessarily also nonsensical. Familiar, human, but nonsensical — consider how often a human disagreeing with another will use the phrase "that's just common sense!" I think the larger models are consuming in the order of 100k as much as we do, and while they have a much broader range of knowledge, it's not 100k as much breadth.
- steveBK123 2y agoWell it's a breadth & depth problem isn't it? Humans are nonsensical, but in somewhat predictable error rates by domain, per individual. So you hire people with the skillsets, domain expertise, and error rates you need. With an LLM, it's nonsensical in a completely random way prompt to prompt. It's sort of like talking into a telephone and sometimes Einstein is on the other end, and sometimes it's a drunken child. You have no idea when you pick up the phone which way its going to go. We feed these things nearly the entirety of human knowledge, and the output still feels rather random. LLMs have all that information and then still have a ~10% chance of messing up simple mathematical comparison that an average 12 year old would not. Other times we delegate much more complex tasks to LLMs and they work great! But given the nondeterminism it becomes hard to delegate tasks you can't check the work of, if it is important.
- phreeza 2y agoI haven't worked with LLMs enough to know this but I wonder: are they nonsensical in a truly random way or are they just nonsensical on a different axis in task space than normal humans, and we perhaps just haven't fully internalized what that axis is?
- steveBK123 2y agoI'm not really sure, and you can pull lots of funny examples where various models have progress & regressions dealing with such mundane simple math. As recently as August "11.10 or 11.9 which is bigger" came up with the wrong answer on ChatGPT and was followed with lots of wrong justification for the wrong answer. Even follow up math question "what is 11.10 - 11.9" gave me the answer "11.10 - 11.9 equals 0.2" We can quibble about what model I was using, or what edge case I hit, or how quick they fixed it.. but this is 2 years into the very public LLM hype wave so at some point I expect better. It gives me pause in asking more complex math questions I cannot immediately verify results, in which case, again why would I pay for a tool to ask questions I already know the answer to?
- jewelry 2y agoThis error is not nonsensical though as normal elementary kids would make similar error and with good episodic memory the agent will fix itself.
- ben_w 2y agoHe did say "sometimes Einstein is on the other end, and sometimes it's a drunken child. You have no idea when you pick up the phone which way its going to go.", so I think that's still a valid thing for him to complain about. LLMs totally violate our expectations for computers, by being a bit forgetful and bad at maths.
- steveBK123 2y agoYes, to put a point on it - How many dollars per month would someone be willing to spend for a chatbot that has a 3rd graders ability at math? Personally, $0 for me. But what if it's a PHD Math degrees ability at math? Tons, in some applications it could be worth $100s or $1000s in an enterprise license setting. But what if it's unpredictably, imperceptibly question to question, 95% PHD and 5% 3rd grader? Again, for me - $0. (not 95% of $1000s, but truly, $0)
- heresie-dabord 2y ago> Humans on the other hand have developed a more elaborate scheme to process, or reason [ ... ] We listen to some explanations, a YT video, a few exercises Frequent repetition in the sociological context has been the learning technique for our species. To paraphrase Feynman, learning is transferring.