4 ms·
I would have said d. XXX. My reasoning being that in sets 1 and 2, the third tile has an X in every space where the first or second tile has an X. Anyway, it
by welshwelsh 3y ago
I would have said d. XXX. My reasoning being that in sets 1 and 2, the third tile has an X in every space where the first or second tile has an X.
Anyway, it seems to me that you are holding AI to a higher standard than humans.
Abstract reasoning abilities do not imply the ability to solve every abstract reasoning problem. Even humans cannot do that. If there is even a single abstract reasoning problem ChatGPT can solve, that means it has abstract reasoning abilities.
Similarly, general intelligence does not imply the ability to solve any problem. It means problem-solving abilities that are not restricted to a specific task or domain. ChatGPT has general intelligence because it can handle situations that it was not explicitly programmed to handle in a wide variety of contexts.
- merlincorey 3y ago> I would have said d. XXX. My reasoning being that in sets 1 and 2, the third tile has an X in every space where the first or second tile has an X. This is of course the correct reasoning, but ChatGPT was way off, as we'll explore shortly. It's just predicting the next likely token and while that works for a lot of things, clearly, it does not seemingly work well for abstract reasoning and logic. While I think it's fair to say that being able to solve one Abstract Reasoning problem would indicate having Abstract Reasoning abilities to some degree, given that it's a multiple choice question there is a solid chance of just lucking into it (like with any human taking a test, even!). As such, these tests typically have multiple questions of the same form but with different specific instances (much like the abstract Class vs the concrete Instance in Object Oriented Programming) in order to identify whether or not a conscious being does indeed have Abstract Reasoning skills. So I will be more than happy to see ChatGPT solve several of these types of problems in a later iteration if it can, but I do think we've found a limitation of the class of models here. Let's look at ChatGPT's reasoning for why that could be: > 1st column: In both sets, the first tile is “O”, so no filling pattern can be determined. False; both full sets have an X in the first column in the second row so this is completely wrong. > 2nd column: In both sets, the second tile alternates between “X” and “O”. False; only the second set alternates between "X" and "O" while the first set has two consecutive "O"'s in the first tile > 3rd column: In both sets, the third tile repeats the same filling pattern as the second tile. False; only the second set repeats the same filling pattern as the second second tile while the first set clearly is a combination of disparate tiles (I specifically chose that representation to make clear that it was an Additive rather than Subtractive problem, by the way) All three of these observations are incorrect but stated authoritatively that many people might not even have noticed. Of course, the "predict next token" engine doesn't KNOW that so it just keeps on predicting tokens into the wrong solution. I believe it will do this most of the time and that shows it does not yet have Abstract Reasoning skills.