4 ms·
> 5. A student might complain about a math exam requiring integration or differentiation by hand, even though math software can produce the correct answer insta
by thomasahle 1y ago
> 5. A student might complain about a math exam requiring integration or differentiation by hand, even though math software can produce the correct answer instantly. The teacher’s goal in assigning the problem, though, isn’t finding the answer to that question (presumably the teacher already know the answer), but to assess the student’s conceptual understanding. Do LLM’s conceptually understand Hanoi? That’s what the Apple team was getting at. (Can LLMs download the right code? Sure. But downloading code without conceptual understanding is of less help in the case of new problems, dynamically changing environments, and so on.)
Why is he talking about "downloading" code? The LLMs can easily "write" out out the code themselves.
If the student wrote a software program for general differentiation during the exam, they obviously would have a great conceptual understanding.
- autobodie 1y agoIf the student could reference notes a fraction of the size of the LLM then I would not be convinced.
- exe34 1y agoI suspect human memory consists of a lot more bits than LLMs encode.
- autobodie 1y agoI rest my case — the question concerns a quality, not a quantity. These juvenile comparisons are mere excuses.
- exe34 1y agoOh we've shifted the goal post to quality now, very good! That does rest the case.
- thomasahle 1y agoExactly. If the paper title had been "LLMs are not that great at thinking", nobody would have had an issue.
- exe34 1y agoI trust you'd have come up with something.
- Workaccount2 1y agoLLMs are (suspected) a few TB in size. Gemma 2 27B, one of the top ranked open source models, is ~60GB in size. LLama 405B is about 1TB. Mind you that they train on likely exabytes of data. That alone should be a strong indication that there is a lot more than memory going on here.
- sigotirandolas 1y agoI'm not convinced by this argument. You can fit a bunch of books covering up to MSc level maths on less than 100MB. After that point, more books will mostly be redundant information so it doesn't need much more space for maths beyond that. Similarly TBs of Twitter/Reddit/HN add near zero new information per comment. If anything you can fit an enormous amount of information in 1MB - we just don't need to do it because storage is cheap.
- Workaccount2 1y agoPeople aren't claiming that they are holding textbooks in their model, that would just be even more evidence of reasoning (the LLM would have to reason what textbook to reference, and then extrapolate from the textbook(s) how to solve the problem at hand - pretty much what students in school do; study the textbook and reason from it to answer new test questions) People are claiming that the models sit on a vast archive of every answer to every question. i.e. when you ask it 92384 x 333243 = ?, the model is just pulling from where it has seen that before. Anything else would necessitate some level of reasoning. Also in my own experience, people are stunned when they learn that the models are not exabytes in size.
- sigotirandolas 1y agoI think the more realistic argument is that the model can generalize, but only by learning shortcuts (e.g. how to pattern match a problem to a likely answer) and simple algorithms (e.g. how to propagate carries in a multiplication). And this only appears intelligent because this pattern matching is really good and backed by a huge amount of compressed/memorized answers. The AI pessimist's argument is that there's a huge gap between the compute required for this pattern matching, and the compute required for human level reasoning, so AGI isn't coming anytime soon.