4 ms·
Why the Abstraction and Reasoning Corpus is interesting and important for AI
- yamrzou 4y agoRelated: ARC 2 — https://news.ycombinator.com/item?id=32740353 https://news.ycombinator.com/item?id=32740353 The Measure of Intelligence — https://news.ycombinator.com/item?id=21547958 https://news.ycombinator.com/item?id=21547958 Neural Abstract Reasoner — https://news.ycombinator.com/item?id=25167182 https://news.ycombinator.com/item?id=25167182 A first lesson in meta-rationality (Bongard problems) — https://news.ycombinator.com/item?id=27411960 https://news.ycombinator.com/item?id=27411960 A HN dicussion about ARC — https://news.ycombinator.com/item?id=29867200 https://news.ycombinator.com/item?id=29867200
- extr 4y agoWhy on earth would you tell the model the task in english? The tasks have extremely simple JSON representations that would make for a way, way more natural input. The colorful grids are just to make it easier (for humans) to visualize the symmetries needed. Eg: Here is an (abbreviated) train/test example for task 007bbfb7 from https://github.com/fchollet/ARC/blob/master/data/training/007bbfb7.json https://github.com/fchollet/ARC/blob/master/data/training/00... "train": [{ "input": [[0, 7, 7], [7, 7, 7], [0, 7, 7]], "output": [[0, 0, 0, 0, 7, 7, 0, 7, 7], [0, 0, 0, 7, 7, 7, 7, 7, 7], [0, 0, 0, 0, 7, 7... "test": [{ "input": [[7, 0, 7], [7, 0, 7], [7, 7, 0]], "output": [[7, 0, 7, 0, 0, 0, 7, 0, 7], [7, 0, 7, 0, 0, 0, 7, 0, 7], [7, 7, 0, 0, 0... I imagine it would be pretty easy to fine tune a davinci model on these and then test, no? Has anyone done it?
- nabla9 4y ago> imagine it would be pretty easy to fine tune a davinci model on these and then test, no? Has anyone done it? That was the a competition. Nobody scored well.
- extr 4y agoThe competition was three years ago though.
- taneq 4y agoExactly! The othello research did basically this, and a few people have fed chess games in using the standard notation.
- meltyness 4y agoI would argue maybe that the author's premise depends on some faulty prompts. With some gentle rephrasing, ChatGPT correctly interprets and manipulates the stack problem. I've observed ChatGPT may have some issues with non-standard sentence structure involving, especially with colons. > Imagine a stack of items (from bottom to top) with a cat, laptop, television, and apple. The apple is moved between the laptop and the television. Which item is on top of the stack afterwards? The author seems also to have misinterpreted that Stable Diffusion is the same sequence-sensitive model as Transformer LLMs, which it is not. While it is generative it has minimal capacity for abstraction, since the focus would rather be minimizing the "kernel" representation of each word in the weights. DALL-E has a bit more luck, but requires much more verbose description to get anything like a screen on top of a cat, but that's presumably more due to the bias in the data set of things being displayed on screens, or cats getting on top of things. Something like > A cartoon image of a television above a cartoon image of a cat https://imgur.com/q9NMY0W https://imgur.com/q9NMY0W Illustrates that it's preference rather than capability.
- civilized 4y agoThe author's prompts are all perfectly understandable to humans and expressed in a direct and clear way, and the AIs can't understand them. This comment essentially says "the prompt is faulty because the AI didn't answer it correctly". This bakes the correctness of the AI into the tautology that it always answers "good" prompts, which are defined as those it answers correctly.
- vba616 4y agoWell, humans don't know how to determine if a proposition is correct. So it's no surprise we didn't teach AIs yet! Edit: I don't understand why experts say it failed when they "asked it to reason". It doesn't have pictures of TVs on cats to produce variations of, right? How does "reason" enter into it?
- civilized 4y agoThere are different perspectives on how much intelligence these foundation models have. Some people think they just elaborate on or mash up their training data. Others think they have internal representations that allow for more human-like creativity. Maybe they learn from the examples what cat and TV and on top of mean, and can put those together in a way not seen in the training data.
- civilized 4y agoThe author mentions that Chollet's ARC tasks contain 200 unreleased "hidden" problems. This set off a light bulb for me: hidden problems are going to be increasingly important for benchmarking AI meaningfully. With RLHF, ChatGPT is memorizing the answer to every problem thrown at it. Eventually it may become impossible to test with any public question bank. We are going to go through a period of fooling ourselves, as we always do, then we're going to hit a wall where models are preforming perfectly on benchmarks while still clearly lacking in genuine intelligence. And then we are going to need to test with a concern for security as serious as that used for the gaokao or SAT. Benchmarking will be a paid service performed under highly controlled conditions.
- taneq 4y ago> then we're going to hit a wall where models are preforming perfectly on benchmarks while still clearly lacking in genuine intelligence More like “and then, as we always do, we’re going to redefine ‘intelligence’ to exclude whatever we don’t want to accept as intelligent.”
- civilized 4y agoNo sane person ever signed up for "100% on so and so benchmarks means AGI". That's cargo cult science.
- vba616 4y ago"Consider, for example, the “simple” concept on top of." This is from fifty years ago: Person: Pick up a big red block. Computer: OK. Person: Grasp the pyramid. Computer: I DON'T UNDERSTAND WHICH PYRAMID YOU MEAN. Person (changing their mind): Find a block which is taller than the one you are holding and put it into the box. Computer: BY "IT", I ASSUME YOU MEAN THE BLOCK WHICH IS TALLER THAN THE ONE I AM HOLDING. Computer: OK. Person: What does the box contain? Computer: THE BLUE PYRAMID AND THE BLUE BLOCK. Person: What is the pyramid supported by? Computer: THE BOX. Person: How many blocks are not in the box? Computer: FOUR OF THEM. Person: Is at least one of them narrower than the one which i told you to pick up? Computer: YES, THE RED CUBE. Person: Is it supported? Computer: YES, BY THE TABLE. Person: Can the table pick up blocks? Computer: NO. Person: Can a pyramid be supported by a block? Computer: YES. Person: Can a pyramid support a pyramid? Computer: I DON'T KNOW. Person: Stack up two pyramids. Computer: I CAN'T. Person: Will you please stack up both of the red blocks and either a green cube or a pyramid? Computer: OK. Person: Which cube is sitting on the table? Computer: THE LARGE GREEN ONE WHICH SUPPORTS THE RED PYRAMID. Person: Is there a large block behind a pyramid? Computer: YES, THREE OF THEM: A LARGE RED ONE, A LARGE GREEN CUBE, AND THE BLUE ONE. P.S. "Any sufficiently advanced technology is indistinguishable from a rigged demo"