2 ms·
The reason it is an indictment of their reasoning capability is that—no matter how much energy is spent trying to say they are not—these are really stochastic p
by abernard1 2y ago
The reason it is an indictment of their reasoning capability is that—no matter how much energy is spent trying to say they are not—these are really stochastic parrots: they do not understand the symbols they are operating with. They operate below that level.
The fact they can't operate on full symbols reliably but require sub-symbols via tokens is concrete proof of that. They may add heuristics or build more CoT sub-chains to get around some of these trickier issues later, but this is the state of affairs right now.
All efforts so far require exponential increases in training size to receive logarithmic increases (at best) in accuracy. And now with o1, it requires exponential compute at inference to scale with that sub-logarithmic accuracy.
People have a short memory these days, but around GPT-3, the majority of people on HN and tech "luminary" founders were saying that these would actually have exponential output and diverge. They were wrong. These models are quickly converging to a training set because they are and always were a curve fit. And even there, they are notoriously unreliable for use cases without a human in the loop, because of the intrinsic amount of information entropy that can be packed into the size of these models. But there is nothing mysterious about them.