4 ms·
The article itself is very assertive and makes a lot of generalizations, but if you look at the source of their claims [1] you see that in the first study they
by epups 3y ago
The article itself is very assertive and makes a lot of generalizations, but if you look at the source of their claims [1] you see that in the first study they are using GPT-3.5 and achieve only a 5% score on a reasoning test that largely relies on spatial intuition - some boxes need to be stacked and unstacked sequentially in a convoluted task. Then, they get criticised so they come up with another paper in which they use GPT-4, which has an improved performance of 30% - a 6-fold increase, which the author describes as "modest". They then decide to change the test to a much harder and more convoluted version, where (surprise!) performance drops down once again [2].
I would also have welcomed a comparison to humans. If we apply this test to 100 humans, can we conclude humans don't reason if only 30 get it right?
[1] https://arxiv.org/abs/2206.10498 https://arxiv.org/abs/2206.10498
[2] https://arxiv.org/abs/2305.15771 https://arxiv.org/abs/2305.15771
- 3abiton 3y agoInteresting take. I feel in this discussion, many people are approaching it from the theoretical limitations of LLMs, and you seem the only one taking the experimental approach. Funny enough, many who doubted LLMs capabilities 4 years ago, have come around their emergent capabilities, yet with much skepticism, simply because we still don't understand how these emergent abilties work. I haven't seen a paper comparing the threshold of performance with the LLMs increased capabilities, and what parameter (and their weights) come into play to influence the performance.
- ryanklee 3y ago> would also have welcomed a comparison to humans. Much of the criticism and skepticism around LLMs rests on a double-standard that itself rests on an almost embarrassing lack of understanding of how humans themselves operate.
- xkcd1963 3y agoAbsolutely!