3 ms·
https://arxiv.org/pdf/2311.09247.pdf https://arxiv.org/pdf/2311.09247.pdf
by recitedropper 3y ago
https://arxiv.org/pdf/2311.09247.pdf https://arxiv.org/pdf/2311.09247.pdf
- famouswaffles 3y agoGPT-4 doesn't fail anything here, it's just worse than humans. Maybe I was unclear but I'm not claiming that GPT-4 is as good as the median human in any task.
- recitedropper 3y agoI appreciate the clarification, although I was linking this as clearly ConceptARC has a number of tasks where GPT-4 fails badly compared to humans. On "Extend To Boundary" category GPT-4 scores 0.2 and humans score 0.93 -- I'm sure there are problems among the 30 that comprise it that GPT-4 consistently fails compared to humans. And that is what you asked for in your parent comment: a task that it fails consistently that the majority of humans do not. Claiming GPT-4 doesn't "fail anything" here is a little pedantic, if you are saying saying well hey it doesn't get a 0 score on anything. The conclusion of the paper is literally GPT-4 fails "to robustly form abstractions and reason about basic core concepts in contexts not previously seen in its training data". And yes, I know the whole "generalization is always a data problem" / "humans come with millions of years of training data" take, but if you follow ARC, it's pretty obvious this type of abstraction forming is the clearest spot where models consistently fail and humans excel. Which to me implies less of a data issue and more of a architectural difference that has yet to be overcome.
- famouswaffles 3y agoLate reply but, LLMs are not bad at abstraction in general, just spatially focused abstraction. https://arxiv.org/abs/2212.09196 https://arxiv.org/abs/2212.09196 Also, how the data is presented matters a lot. Currently, LLMs handle the benchmark much better presented linearly in 1 dimension https://arxiv.org/abs/2305.18354 https://arxiv.org/abs/2305.18354