3 ms·
It's not about giving the same answer to the same question. It's about getting the right answer 100% of the time, in some very specific domains. If you know, un
by vhantz 2y ago
It's not about giving the same answer to the same question. It's about getting the right answer 100% of the time, in some very specific domains. If you know, understand and are able to use the basic rules of arithmetic, 2 + 2 only has one answer. If you know, understand and are able to use the basic rules of formal logic, the same premises will lead you to the same conclusion. Two trivial cases for any reasoning person . Two cases that illustrate how fundamentally different LLMs text generation is to reasoning. Two cases that illustrate some of the challenges that need to be solved to bring AI models closer the fiction so many on this site are desperately taking them to be.
Of course those who don't care about improving those systems also don't care about understanding their limits, which is unsurprisingly the case for a lot of people on this website.
- CamperBob2 2y agoYou've failed to explain -- or to understand -- how the models get the right answer at all. The fact is, when you ask what 2+2 is, or what 2342+33222 is, the current ChatGPT model will give you the correct answer, even if you don't tell it to write code to get it. The first answer can simply be regurgitated. The second one, not so much. Heck, let's throw in a square root for the fun of it: https://i.imgur.com/Q9eHAaI.png https://i.imgur.com/Q9eHAaI.png How'd it do that, if it can't reason? That problem wasn't in its training corpus. Similar ones were, with different numbers, and that was enough. Ask it 100 times, and it will probably get it wrong a certain percentage of the time... just like you would if I asked you to perform the calculation in your head. Notice that the model actually got the LSD slightly wrong in this example. 188.58 would be a better estimate. It even screws up the way we do. That, to me, is almost as interesting as the fact that it can deal with the problem at all. Of course those who don't care about improving those systems also don't care about understanding their limits The people who do care about improving these systems seem to be doing a pretty awesome job. As for the limits, they frankly don't seem to exist. They certainly aren't where you and your predecessors over the past few years have assured us they are.