4 ms·
Because it's not inherent to the system itself. For example there's nothing inherent to token prediction to make it capable of doing arithmetic and you wouldn'
by nivenkos 3y ago
Because it's not inherent to the system itself.
For example there's nothing inherent to token prediction to make it capable of doing arithmetic and you wouldn't expect a small/briefly trained model to manage it. But if learning a very large model on a huge dataset leads to it "learning" the decimal number system and arithmetic via token prediction then that is an emergent capability.
It's just like the Chinese Room thought experiment - https://en.wikipedia.org/wiki/Chinese_room#Chinese_room_thought_experiment https://en.wikipedia.org/wiki/Chinese_room#Chinese_room_thou...
- tlb 3y agoYes, in some sense the recent LLM results show us how big the Chinese room has to be: enough for 10^11 parameters. If the index cards Penrose imagined could hold one column of a matrix, it's around 10^8 cards. It's not surprising that people had poor intuitions for what was possible with that level of complexity.
- fenomas 3y ago> I believe that in about fifty years’ time it will be possible to programme computers, with a storage capacity of about 10^9, to make them play the imitation game so well that.. Turing wrote that in 1950, so I suppose his intuition was within a couple of orders of magnitude...
- deleted 3y ago[deleted]
- zarzavat 3y agoPerhaps a bad example although I agree with the general argument. There are large datasets of calculations specifically for training large language models. It’s not just picking it up from reading books. And even then these models suck at calculating, half the time they just make an answer up. Calculation appears to be something that is actually not an emergent property of constant-time token prediction. Which we already knew from Turing anyway.
- Traubenfuchs 3y ago> For example there's nothing inherent to token prediction to make it capable of doing arithmetic ...aren't we just witnessing proof that this is wrong? Apparently those abilities are inherent to token prediction models with sufficient parameter size.
- cowl 3y agoin fact their example of https://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/modified_arithmetic https://github.com/google/BIG-bench/tree/main/bigbench/bench... showed 0 emergent behavior in understating modified arithmetic even in large LLMs. the accuracy (and lack of it) in unmodified arithmetic is a simple token replacement heuristic that comes naturally from the transformer model of "attention".
- golol 3y agoThis is wrong. I asked GPT-4 right now to perform this task and it got 3/3 3-digit calculations correct, and on a 4-digit calculation was off by 10 (not 1!). And note that artihmetic is seen as a weak spot of LLMs, it is not a good example to attack the claim that they have emergent properties.
- nivenkos 3y agoNice find, I think we'll get there eventually though, as models can hold more state.
- famouswaffles 3y agoWe are already there. 100% accuracy on up to 13 digit addition can be taught to 3.5 as is. https://arxiv.org/abs/2211.09066 https://arxiv.org/abs/2211.09066 And 4 has little need for such out the box
- cowl 3y ago> in this work, we identify and study four key stages for successfully teaching algorithmic reasoning to LLMs: (1) formulating algorithms as skills, (2) teaching multiple skills simultaneously (skill accumulation), (3) teaching how to combine skills (skill composition) and (4) teaching how to use skills as tools. So it's not an emergent property of LLM but 4 new capability trainings. Noone is saying you can't teach these things to an Agent, just that these are not emergent abilities of the LLM training. by default a LLM can only match token proximity, all trainings of LLMs improve the proximity matching (clustering) of token but they do not teach algorithmic reasoning. it needs to get bolted on as addon.