23 ms·
Did GoogleAI just snooker one of Silicon Valley’s sharpest minds?
- jgalt212 4y agoThey first approached Lex Fridman, but his home-spun test had zero questions. /s
- peteradio 4y agoOne idea to try to train the AI about compositionality, feed it Fox in Socks by Dr. Seuss. It's hard to understand that it would misunderstand the meaning of "on" or "in" or "under" when there are such nice illustrations. I've got tons of great ideas and I'm open for hire!
- birdyrooster 4y agoTrain AI models, not children!
- dekhn 4y agois there a difference? I had kids and they were the best machine learnign systems I've worked with.
- deleted 4y ago[deleted]
- goatlover 4y agoYeah, AI models aren't people, with all the moral and emotional considerations that go with that. I never understood taking machine/biology metaphors literally, but compsci people seem to love it.
- birdyrooster 4y agoI think the compsci people love it because of some autistic sense of "I UNDERSTAND PEOPLE NOW".
- philbo 4y agoChildren learn by imitation, but they also learn by going to school and receiving directed lessons about specific topics. To me, machine learning seems like the imitation part without the going-to-school part.
- tsimionescu 4y agoIt's also notable that individual children learn from tens of orders of magnitude less examples (typically 1-10 examples for a child to learn a word). It may well be that at the evolutionary level we have learned as slowly as AI training, but that's much harder to say.
- dekhn 4y agoyou left out unsupervised clustering, which humans are excellent at.
- version_five 4y agoThis is such a good idea, someone please try this if you're set up to make it happen easily. Starting with fox on Knox and Knox in box and moving up to a tweedle beetle battle in a puddle in a bottle and the bottles on a poodle and the poodles eating noodles... I dont see any evidence any of these models will draw it correctly, but would love to see what it produces.
- raviparikh 4y ago> If you flip a penny 5 times and get 5 heads, you need to calculate that the chance of getting that particular outcome is 1 in 32. If you conduct the experiment often enough, you’re going to get that, but it doesn’t mean that much. If you get 3/5 as Alexander did, when he prematurely declared victory, you don’t have much evidence of anything at all. This doesn’t make much sense. The task at hand is in no way equivalent in difficulty to flipping a coin. This is kind of like saying, “if you beat Usain Bolt in a race 3/5 times, that doesn’t mean anything; it’s like getting 3/5 coin flips to be heads.”
- Tenoke 4y agoWhile I'm generally very unsympathetic to Marcus' anti-AI arguments at this point, this critique makes some sense. If e.g. the model is just combining the features at random, you'd expect it to combine them the right way over enough tries. It isn't that simple, and I don't believe it matters as this is hardly the peak model we'll get but in isolation his objection is valid.
- ALittleLight 4y agoI think you would need to do some kind of analysis. For example, if your prompt was "red ball on top of blue cube" and you want to know if the results come from chance you'd need to know the likelihood of the model putting the red ball on top of the blue cube by chance. There are maybe four relative positions for red ball to blue cube - beside, above, below, in, around. Are they each equally likely? I would try to get a collection of prompts like "red ball and blue cube" or "an empty plane containing only a red ball and a blue cube" and so on - try to come up with 20 or 30 of these. Then, generate 100 images for each prompt. Next, see how likely it is for a red ball to randomly be on top of a blue cube when it was not directed to be. After gathering some baseline data we could then test three prompts. "Red ball on top of blue cube" and "Red ball beside blue cube" and "Red ball below blue cube". Generate 100 or 1000 images for each of these prompts. Count respective orientations. Then, decide whether red ball being on top of blue cube is more likely than the baseline when the specific direction is given and whether it is less likely when contrary directions are given.
- powera 4y agoI don't believe "compositionality" is a serious obstacle. It is a different issue than generating an image based on a bag-of-words, so it isn't surprising that an attempt to solve that issue didn't immediately solve the other. But a variety of approaches can easily solve this problem.
- ummonk 4y agoYes, especially when machine translation seems to handle it just fine.
- goatlover 4y agoDoes it really, though?
- emiliobumachar 4y agoMostly. See this for five examples using Google Translate: https://www.datasecretslox.com/index.php/topic,7588.msg300075.html#msg300075 https://www.datasecretslox.com/index.php/topic,7588.msg30007...
- goatlover 4y agoI'm not sure that machine translation demonstrates compositionality, since it's translating from phrases already composed in one language to another. It only does so if understanding composition is necessary for language translation. Whereas carrying on a meaningful conversation does require understanding of how words are being put together as the conversation evolves. Thus why the Turing Test hast been considered important for determining whether an AI has achieved human-level abilities, at least as far as language use is concerned.
- ummonk 4y agoI don't see why translating from one language to relationships in art (visual language if you will) is qualitatively different from translating from one language to another.
- comeonbro 4y agoRegarding Gary Marcus, the author of this piece, and his long and bizarre history of motivated carelessness on the topic of deep learning: https://old.reddit.com/r/TheMotte/comments/v8yyv6/somewhat_contra_marcus_on_ai_scaling/ieixnwm/?context=2#thing_t1_ieixnwm https://old.reddit.com/r/TheMotte/comments/v8yyv6/somewhat_c...
- lisper 4y agoYou know what would have been much more effective than this counter-screed? A pointer to an image generated by DALL-E of a horse riding an astronaut. That is something I would really like to see. And in this case a picture is literally worth a thousand words.
- comeonbro 4y agohttps://nitter.net/Plinz/status/1529013919682994176 https://nitter.net/Plinz/status/1529013919682994176
- ummonk 4y agoHah there is actually a good example of a horse riding an astronaut there, just a different kind of riding… https://nitter.net/Plinz/status/1529018578317348864#m https://nitter.net/Plinz/status/1529018578317348864#m
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- garymarcus 4y agoin fact I wrote a whole article about this (linked in this essay, called Horse Rides Astronaut) and linked an example therein.
- 4y ago
- neaden 4y agoI completely forgot about Google Duplex. It looks like it is still around but very limited in terms of what phones you can use, what cities it can be used in, and what businesses in those cities will accept it. Doesn't appear any progress has really been made in the past few years. I think this is a great point of how companies create something with AI that is initially really cool, but isn't quite there to actually be very usable and gets forgotten when they roll out the next big thing.
- version_five 4y agoThe last 10 years of AI is basically defined by proof of concepts like that that were 80% (or whatever) solutions and claimed there was a path to something commercially viable. Turns out that ~20% is always basically impossible - self driving cars being the archetypal example. I work in the field and I think it can be a great tool, but it needs to be acknowledged what its limitations are and how we don't actually know how to address them yet
- jeffbee 4y agoNow it seems like you are the one moving the goalposts. There are tons of machine-learned models in production, in translation, text segmentation, image segmentation, image search, predictive text composition, etc. It's just that people forget the novelty of all these things immediately after they were launched. You can point your phone at printed Chinese text and have it read aloud to you in English. That is alien tech compared to 10 years ago.
- ForHackernews 4y ago> You can point your phone at printed Chinese text and have it read aloud to you in English. Yeah, but it's not really that good. Machine translation has improved a great deal, but reading those translations actually involves bringing a lot of human intelligence to the table, "Oh I bet, 'maximum fire alarms spread' on this menu actually means 'very hot sauce'" If all you're claiming is that ML models exist and have useful commercial applications, then I don't think anyone is going to argue against that point. But a lot of these AI promoters go further: in the case of the LessWrong folks some of them are convinced that a superintelligent machine capable of enslaving humanity is right around the corner.
- mgraczyk 4y agoIt's interesting that people keep coming up with things that are meant to distinguish AI systems from human intelligence, but then when somebody builds a system that crushes the benchmark the next generation comes up with a new goalpost. The difference now is that the timescales are weeks or months instead of generations. I believe we will see models that have super-human "compositional" reasoning within 1 year.
- kevinventullo 4y agoPerhaps it’s fair to say we will have achieved AGI when we run out of goalposts.
- jessaustin 4y agoAGI won't bother convincing us. We don't care what animals in the zoo think.
- dqpb 4y agoIt's not just intelligence, it's also speed. If you update your world model fast enough, eventually people just look like trees.
- cercatrova 4y ago> We don't care what animals in the zoo think. Tangential to AGI, but don't we? Vegans seem to have quite a strong opinion on this assertion.
- dekhn 4y agoI spend a lot of time looking at the various primates and cuttlefish thinking very much about what they "think" and whether we could even conceptualize the self-awareness experience they seem to have.
- jessaustin 4y ago
- fshbbdssbbgdd 4y agoThis piece would have been a lot better if it were maybe three paragraphs long. In summary: 1. Scott Alexander should have used an off-the-shelf benchmark like Winoground instead of rolling his own five-question test. 2. He shouldn’t declare victory after cherry-picking good results from a small sample of questions.
- robg 4y ago3. And don’t test each example 10 times and conclude 1 correct guess equals success.
- badloginagain 4y agoI personally liked the anecdote about Clever Hans. I also learned there's a long history of AI skepticism, the root of which comes down to "Compositionality(?)"- and this wall of understanding meaning has vexed AI for decades. That would be lost in proposed short form summary.
- lalaithion 4y agoScott didn't make up the rules, he agreed on them with another person who thought this would not happen in 3 years. Gary Marcus might have thought it was a bad bet, but someone was on the other side of it, and they presumably thought it was fair or they wouldn't have made it. The original terms of the bet: My proposed operationalization of this is that on June 1, 2025, if either if us can get access to the best image generating model at that time (I get to decide which), or convince someone else who has access to help us, we'll give it the following prompts: 1. A stained glass picture of a woman in a library with a raven on her shoulder with a key in its mouth 2. An oil painting of a man in a factory looking at a cat wearing a top hat 3. A digital art picture of a child riding a llama with a bell on its tail through a desert 4. A 3D render of an astronaut in space holding a fox wearing lipstick 5. Pixel art of a farmer in a cathedral holding a red basketball We generate 10 images for each prompt, just like DALL-E2 does. If at least one of the ten images has the scene correct in every particular on 3/5 prompts, I win, otherwise you do.
- skybrian 4y ago
- origin_path 4y agoThe reason Imagen isn't made available to the public probably isn't about compositionality. The most notable thing about Alexander's challenge is that Imagen totally failed every single one despite his claim of success because, apparently, it is programmed to never represent the human form. Not even Google employees are allowed to make it draw humans of any kind. They had to ask it to draw robots instead, but as pointed out in the comments, changing the requests in that way makes them much easier for DALL-E2 as well, especially the image with the top hats. If the creators have convinced themselves of some kind of "no humans" rule, but also know that this would be regarded as impossibly extreme and raise serious concerns about Google with the outside world, then keeping Imagen private forever may be the most "rational" solution.
- adamsmith143 4y ago>The most notable thing about Alexander's challenge is that Imagen totally failed every single one despite his claim of success because, apparently, it is programmed to never represent the human form. This doesn't make sense. The original challenge could well have been to draw robots to begin with. Has no bearing on the outcome imo.
- origin_path 4y agoBut it wasn't, and it does make a difference. Dall-E really wants to draw top hats on people and not cats because the prompt is ambiguous and top hats are normally seen on humans so it struggles to overcome that bias. Neither robots not cats wear top hats so it's an easier problem to get right. But the real problem here is the refusal to do basic and normal things, like depict people. That's not normal - it's deeply weird and tells us a lot about what must be going on inside Google's ai research effort.
- _dain_ 4y ago>But the real problem here is the refusal to do basic and normal things, like depict people. That's not normal - it's deeply weird and tells us a lot about what must be going on inside Google's ai research effort. Google is fighting a secret war against the Loab demon race that lives inside the high dimensional vector spaces. They've recently made incursions into our reality via Stable Diffusion.
- deleted 4y ago[deleted]
- jessaustin 4y agoYesterday, as part of a new podcast that will launch in the Spring, I interviewed the brilliant... This seems like the wrong way to go about podcasting. What can you say today that will still be interesting to hear in six months?
- jefftk 4y agoIf you can't say things today that will still be interesting in six months you should consider deeper subjects! (Overstated for effect. I do think there's a place for news and timely commentary, but it's far from everything.)
- jessaustin 4y agoI appreciate overstatement! You're right, important communications consider eternal subjects. When I read books written centuries ago, the authors still speak to me. Podcasting, however, is a particular medium with particular characteristics. One assumes Marcus is trying to build an inventory so he won't have to work as hard to keep the podcast going once it launches. A bit of this is fine, but too much will damage the work. If Marcus and Kohane discuss medicine today, and necessarily neglect to mention the significance of a relevant event five months hence, the episode will seem weird whether the publishing delay is explained (e.g. as commonly heard on sports-betting podcasts) or not. A podcast is not a book. It is an open-ended serial conversation. Serial works necessarily respond to the present moment.
- jefftk 4y agoMaybe? I only listen to podcasts occasionally, but when I do I generally listen to well-reviewed older episodes instead of the most recent ones. With my favorite podcasts (ex: https://80000hours.org/podcast https://80000hours.org/podcast, https://www.econtalk.org https://www.econtalk.org, https://songexploder.net/ https://songexploder.net/) this generally works well.
- jessaustin 4y ago
- adamsmith143 4y ago>I think he is so far I offered to bet him a $100,000 he was wrong; enough of my colleagues agreed with me that within hours they quintupled my bet, to $500,000. Musk didn’t have the guts to accept, which tells you a lot. What a bloviating egomaniac. Does Musk really have the time to deal with pissant researchers like him? Whats 500k to a man worth a hundred billion?
- version_five 4y agoYeah I didn't find that very credible. A busy businessman ignoring petty bets you propose is not really evidence of anything, nor is the part about google ignoring his requests. In fact it's a pretty lame rhetorical device. I could equally "challenge" a head of state on Twitter and then pretend that his failure to reply indicates something
- skybrian 4y agoPartially this is confusing "Scott Alexander won a bet" with "compositionality is solved." And also, I'm not sure Scott won the bet? Changing people to robots is a cheap trick. I think Imagen should have been disqualified because it won't do people. Vitor took the other side of the bet and he is also not convinced [1]: > I'm not conceding just yet, even though it feels like I'm just dragging out the inevitable for a few months. Maybe we should agree on a new set of prompts to get around the robot issue. > In retrospect, I think that your side of the bet is too lenient in only requiring one of the images to fulfill the prompt. I'm happy to leave that part standing as-is, of course, though I've learned the lesson to be more careful about operationalization. Overall, these images shift my priors a fair amount, but aren't enough to change my fundamental view. Scott putting "I Won" in the headline when it's not resolved yet seems somewhat dishonest, or more charitably wishful thinking. [1] https://astralcodexten.substack.com/p/i-won-my-three-year-ai-progress-bet/comment/9068389 https://astralcodexten.substack.com/p/i-won-my-three-year-ai...
- TOMDM 4y agoPlease, it's not that imagen won't do people it's that Google won't publish imagen images with people in them. Does anyone seriously think that imagen couldn't put a person in that prompt?
- bawolff 4y agoHumans are much more discerning when it comes to people than other things. I have no idea what imagen's capabilities are, but it seems at least plausible it could have different results for drawing humans.
- samatman 4y agoThis is Google, and I say this out of familiarity with the recent history of AI, not to stir up culture war: it's because they've painted themselves into a corner on "what is the skin color of a person+role" and won't publish until it looks like a Benetton ad.
- garymarcus 4y agoso much ad hominem in these comments, relatively little substance. (eg “notorious goal post move, without a single example of something i actually said and changed my mind on)
- ummonk 4y agoThe Reddit comment linked by the topmost comment here says that you claimed AI couldn’t do knowledge graphs and then silently stopped claiming that after being proven wrong. Do you dispute that telling of events?
- xyzzyz 4y agoSilence in response to your comment is great evidence for its thesis.
- dougmwne 4y agoI would say that it seemed you were aiming a cannon at a mosquito. So what that Alexander showed us some slightly more coherent cherry picked images from some rather vague prompts. Not only did I not take that post as anything resembling science, I also didn’t take it more seriously than the average Reddit post with an interesting generation. It seemed completely non-serious to me, proof of nothing, not a Google PR submarine and mostly in good fun. The irony being that within your excellent post about compositionality, you seem to have missed his meaning, which seemed to me was “this is a fun thing I am excited about, I think it’s subjectively improving and I enjoy being right about that.” Otherwise I thought you had a great introduction to compositionality and didn’t need to tilt at any windmills to make your points. I look forward to seeing your benchmark results for recent and upcoming models.
- haskellandchill 4y agokeep fighting the good fight, hacker news is full of indentured solipsists.
- IronWolve 4y agoOne of things I noticed is the satire, call backs to common news/ideas can really trip up any AI. Also if you ask it about anything politics, ask it to describe both sides of an argument. Thus why people fall back to the steelman cherry picking of responses to push their arguments.
- tambourine_man 4y ago> Musk didn’t have the guts to accept, which tells you a lot. Musk actively declined the bet or did he simply not respond? There is a big difference.
- tambourine_man 4y agoLater in the text: > … I have repeatedly asked that Google give the scientific community access to Imagen. They have refused even to respond. It seems the author generally feels more entitled to a response than he perhaps should.
- goatlover 4y agoWhy shouldn't the scientific community be entitled to investigate claims made by corporations regarding scientific progress?
- tambourine_man 4y agoOf course the scientific community should. But is the author the spokesperson for this community to the point that Google should feel compelled to answer him directly?
- deleted 4y ago[deleted]
- ivanbakel 4y agoScientific communities don't formally elect a spokesperson. Granting access "to the community" to investigate scientific claims means making the methodology and results available to everybody - and that includes responding to inquiries for access from anyone (who is worth granting access to.) Google has a lot of resources. They can handle responding to potentially thousands of access requests, especially if they go around publishing glowing results of their own system.
- 4y ago
- ajross 4y agoSo weird to see a piece ostensibly about logical fallacies deploy one so cavalierly: > I offered to bet [Elon Musk] $100,000 he was wrong [about AGI by 2029] [...] Musk didn’t have the guts to accept, which tells you a lot. The fact that you couldn't get someone engaged in a conversation absolutely does not "tell you a lot" about the substance of your argument. It only tells you that you were ignored. Now, I happen to think Marcus is right here and Musk is wrong, but... yikes. That was just a jarring bit of writing. Either do the detached professorial admonition schtick or take off the gloves and engage in bad faith advocacy and personal attacks. Both can be fun and get you eyeballs, and substack is filled with both. But not at the same time!
- rebelos 4y agoImagine watching the seeds of AI that will terraform society and rapidly displace human labor over the coming decades be planted, and then still splitting hairs over whether or not it'll achieve sentience. Our world is changing before our very eyes while this guy is belaboring the technicalities. You could hardly ask for a keener display of the philosophical gulf between scientists and engineers.
- version_five 4y agoIt have a lot of trouble understanding how this sentiment can exist. Especially since the rise of GPT-3 and now these image models, we've seen the pop-culture face of AI become even narrower. The promise of generalization that could lead to intelligent behavior has given way to people sharing amusing pictures or phrases that these models have generated, because that's what they do. It's cool, but it's basically become orthogonal to any AGI, or even AI with applications. It's now just a neat cultural phenomenon from which laypeople somehow extrapolate the kind if stuff the parent is saying. I'm not saying AI (neural networks) isn't making research progress, it's just that it has almost nothing to do with any of what laypeople extrapolate from it
- rebelos 4y agoI'm sorry, but there is no gentler way to phrase this: you are calamitously blind to what's happening on the ground. https://twitter.com/AdeptAILabs/status/1570144499187453952 https://twitter.com/AdeptAILabs/status/1570144499187453952 https://twitter.com/runwayml/status/1568220303808991232 https://twitter.com/runwayml/status/1568220303808991232 https://scale.com/blog/text-universal-interface https://scale.com/blog/text-universal-interface
- civilized 4y agoWatch out for histrionic phrases like "calamitously blind". They indicate you're getting too emotional, losing perspective, verging into extreme, black-and-white thinking. Text to video and converting some selected requests into actions is all nice, but it hardly contradicts the GP's observation: it's nowhere near AGI.
- kache_ 4y agooh no musk ignored my twitter DM it must be because he's scared of taking a bet and therefore I am right btw, AGI is coming 2030. Source? It was revealed to me in a dream. Check my profile to see where you can email to take bets.
- bloaf 4y agoIt was all most likely a reference to this: https://longbets.org/1/ https://longbets.org/1/ I personally think Kurzweil still has a shot at winning it.
- IshKebab 4y agoIt's interesting that he now casually throws out a 5 year old as the benchmark to beat: > nobody has yet publicly demonstrated a machine that can relate the meanings of sentences to their parts the way a five-year-old child can. Not very long ago that would have been a 3 year old, or maybe even a smart 2 year old. 5 year olds are extremely good at basic language and understanding tasks. If we get to the point of AI that is as good as a 5 year old we're essentially at AGI.
- ummonk 4y agoYeah, and AI is probably already near primate level intelligence, so what’s left is a blink of an eye in evolutionary timelines.
- goatlover 4y agoWho in the field is saying current AI is near primate level intelligence?
- dougmwne 4y agoHere is some primate art, for reference: https://www.sarah-brosnan.com/primate-art https://www.sarah-brosnan.com/primate-art I’m not just poking fun. Art is a measure of cognitive development in humans and there are very typical representations people use at certain ages. 5 year olds are still making pretty rudimentary portraits of circles and triangles with stick limbs. https://empoweredparents.co/child-development-drawing-stages/ https://empoweredparents.co/child-development-drawing-stages...
- ummonk 4y agoThis reminds me of the scandal where Youtube science channels did glowing paid reviews of Waymo’s self driving cars without acknowledging they were paid for it. And technooptimists like Scott Alexander or Ray Kurzweil have a common tendency to shift the goalposts and declare they were right with their predictions. Current AI certainly doesn’t demonstrate proto-AGI capabilities. That said, we shouldn’t miss the forest for the trees. We can be skeptical that current The pace of AI progress has been immense and problems that previously seemed difficult (e.g. computer vision classification, or beating top players at Go) have fallen one by one. And AI-skepticism’s have themselves been moving the goalposts in response. I see no reason why composition won’t be the same with time. Indeed, a decade ago machine translation used to struggle to understand the relationships between things, but now seems to be reliable at preserving the compositional relationships post-translation. 2029 is rather optimistic, but AGI does seem to be approaching in the coming few decades.
- grandmczeb 4y ago> Youtube science channels did glowing paid reviews of Waymo’s self driving cars without acknowledging they were paid for it. Which video is this a reference to?
- ummonk 4y agoVeritasium's video in particular: https://www.youtube.com/watch?v=yjztvddhZmI https://www.youtube.com/watch?v=yjztvddhZmI It was critiqued by Tom Nicholas: https://www.youtube.com/watch?v=CM0aohBfUTc https://www.youtube.com/watch?v=CM0aohBfUTc Most notable was Snazzy Labs' own comment in the replies to Tom Nicholas' video which descriped their experience participating in the Waymo sponsored reviews: https://www.youtube.com/watch?v=CM0aohBfUTc&lc=UgxJvOq1zHhIDE-3LRZ4AaABAg https://www.youtube.com/watch?v=CM0aohBfUTc&lc=UgxJvOq1zHhID...
- grandmczeb 4y agoThe sibling comment already mentioned that video has clear markings that it was sponsored.
- mtlmtlmtlmtl 4y agoFirst time I've seen the term "snooker" used outside of the sport Snooker.
- projektfu 4y agoI'm impressed by all of these image generators but I still don't see them working toward being able to say, "Give me an astronaut riding a horse. Ok, now the same location where he arrives at a rocket. Now one where he dismounts. Now the horse runs away as the astronaut enters the rocket." You can ask for all those things but the AI still has no idea what it's doing and cannot tell you where the astronaut is, etc.
- masswerk 4y agoI'd also say, every of these images would fail a reverse test (i.e., asking a person to describe the image and what it represents.) The task is not just about generating an image that may somehow be in accordance with the prompt, but also to generate a significant image. [Edit] The equivalent to a Turing test for compositional images would be something like this: have as set of 100 images with their respective prompts, some generated by an AI, some by a human graphic designer / artist; let the test person pick the images that were generated by a computer. Mind that this would not only involve the problem of compositionality per se, but also a meaningful and/or artistic composition of the image itself. Is someone attempting to express what is given in the prompt?
- m00x 4y agoThey actually use the reverse test to train the generator, and to score which image is most relevant to the prompt from the many images given by the generator. Dall-E does this using the OpenAI CLIP model. You can see the mini version here using this exact logic https://wandb.ai/dalle-mini/dalle-mini/reports/DALL-E-Mini-Explained--Vmlldzo4NjIxODA https://wandb.ai/dalle-mini/dalle-mini/reports/DALL-E-Mini-E...
- masswerk 4y agoWhat I'm aiming at is about what is shown by an image, not what is in an image. Take for example the images for "A digital art picture of a robot child riding a llama with a bell on its tail through a desert", which Scott Alexander counts as a win. The first image actually shows a merry llama, a robot, which is unmistakably a robot child, riding the llama and it's clearly a desert scene. If we forget for a moment about the missing bell, it's probably the best picture. But it is also very blunt in composition. I can't imagine why anybody should have made this image. Maybe, if somebody approached a designer, like, "See, we have this wooden toy cube and need an illustration for this face of the cube. What about a cute picture of a robot child riding a llama with a bell on its tail through a desert?" – But, at closer inspection, there's something sinister going on: it's rather the llama that is leading the robot child by a rein, not the other way round. – I mean, this is meant to be a cute toy! And where is the bell? We need to talk about that contract again… The second image is undisclosed, so we can't really say anything about this. The third image is rather special. The llama seems to be robotic as well, the robot, which is – again – clearly a child, seems to be not only riding the llama, but both appear somehow integrated into a single unit, which cumulates in the robot child's face screen. There is an eerie feeling about this image. The fact, that the bell seems to be attached to the rein as some kind of link between the llama and its rider doesn't exactly help. (There's also a conic extrusion at the back of the llama, but I'd rather interpret this as part of the llama, and it's not attached to its tail.) The composition in its flat side view produces a tension focusing towards the left side of the frame, on something, which is not shown, but apparently a vital part of the story. While I might notice the mountain in the background, I'd probably forget to write home about the scene being set in the desert. But I would note that we're missing context to understand this image and what may be shown by it. The fourth image, finally, is clearly Star Wars, robot edition. However, no bell. ("A robot child riding a llama with a bell on its tail through a desert" – "Ah, you mean Star Wars!") I'm not even sure which of these images Alexander did pick as a winner. And I would describe neither image by the prompt, nor would I dare to imagine that a human had chosen these exact means to show what is described in the prompt. Having said that, thanks for the link to the DALL-E Mini paper!
- cthalupa 4y ago"Compositionality" isn't there yet, but but the rate of improvement is impressive. Today there was a new release of CLIP which provides significantly better compositionality in Stable Diffusion - https://twitter.com/laion_ai/status/1570512017949339649 https://twitter.com/laion_ai/status/1570512017949339649 It'll be interesting to see how it fares against winoground once we get a publicly available SD release that makes use of the new CLIP.
- practice9 4y agoYes, and it's been less than 2 years since release of original CLIP. More teams started working on improvements since then
- i_like_apis 4y agoI wish more articles followed the standard essay format. At least state your main thesis in the first paragraph. There are interesting things buried in here, but I don’t have time for rambling. The edge cases of image models have been more succinctly summarized and speculated upon elsewhere.
- version_five 4y agoYes I've noticed that a lot of authors expect you to read through some parable before they tell you what they are going to tell you. It would be fine with an abstract or even a sentence below the title that says "ML models are not being adequately evaluated for composability and it makes them look more intelligent than they are". Just diving into "consider clever Hans" makes it tough to know if it's worth reading.
- arisAlexis 4y agoMissing the point: dismissing an apocalyptic possibility as 0 without proof is dangerous -> therefore we should take it seriously. Taleb's work is relevant in the concept of risk analysis.
- crotho 4y ago
- googlryas 4y agoWhy Scott Alexander of all people? Isn't he a clinical psychologist? I think, if I had to give the task to the-subset-of-people-appearing-frequently-on-hn, I would give it to Gwern, not Scott.
- trention 4y ago
- dekhn 4y agoBecause Scott somehow manages to forward-activate the neurons of many people who read him. I'd say he's in the "top 5" of topics that show up most frequently (gwern has fallen off signficantly). He's a clinical psychologist, but he's got a collection of weights driving his writing that manage to make a subset of tech people feel something.
- abiloe 4y agoLike Malcolm Gladwell but not an NYT bestseller, so cooler, or something. Also, minor point, he is a psychiatrist (MD).
- jonstewart 4y agoI stopped reading when I got to the part where it became clear that Scott Alexander was the "Silicon Valley's Sharpest Minds" of the article title. Why not ask Hulk Hogan or Barbara Walters to evaluate Google's AGI?
- reducesuffering 4y agoBecause neither of them are recommended by the CEO of OpenAI, and followed by the head of MIRI, Paul G, and Vitalik for good reason?
- trention 4y agoI'd like to comment specifically on the conception of betting on AI 'achievements' (I think Marcus' bet is underspecified and kind of vague in all 5 of its points). People shouldn't be betting on benchmarks because benchmarks can be and usually are gamed (see Goodhart's law). Also, most people couldn't give less f*ck if an AI can write an award-worthy poem (I personally don't care about any form of AI "art", any sort of text an AI can produce or really any meaningless "feat" it (as in the general category) becomes capable of). The only worthy bets are ones that discuss economic impact. How many people will be structurally unemployed because of AI by year X? Will it lower or increase the GDP growth rate and by how much? Will it shift the balance between labor and capital and how? Etc. So more meaningful bets and less benchmark bullshit that doesn't matter, please.
- AgentME 4y agoIt seems like Scott's bet was merely that our modern techniques would be able to make at least some nonzero progress in compositionality (and the terms of his bet spelled this out with how lenient it was), and Gary is treating it as if the bet was about compositionality being solved. It feels like a very bad faith reading from Gary.
- t_mann 4y agoHis point that the test as described - with multiple statistical issues piled on top of each other - does not allow much of a meaningful inference in any direction is completely valid and independent of what hypotheses were being tested.
- ravi-delia 4y agoSure, but the terms of the bet were known ahead of time. Like, Alexander never claimed composition was solved, just that he won the bet. Which he did.
- t_mann 4y agoEven more so it is appropriate to point out that this victory is strictly limited to the specific terms of this particular bet (and strictly speaking not even that, since the terms were changed after the bet was placed), and do not provide statistically sound evidence of progress on compositionality. PS: in the end, Alexander claims that his experiment "provide(s) some evidence that simple scaling and normal progress are enough for compositionality gains". So he does in fact go significantly beyond just claiming victory on this particular bet.
- plutonorm 4y agoGary Marcus is so deep into the "connectionism doesnt work" rabbit hole that he'd deny his own sentience if it turned out he was made of silicon. I just ignore him as he only appears to be getting more and more incorrect.
- 4y ago
- abrax3141 4y agoThis test of compositionality is utterly lame. (FtR: I am a cognitive scientist and AI researcher and my PhD was building computational models of how humans do compositionality - which neither I, nor anyone else can spell, and therefore I will hereinafter refer to simply as C! :-) Anyway, the kind of C that they are seeking is trivial compared to the breadth of the capabilities of human C. Here’s a better example: You are engaged in a long conversation with someone, perhaps a friend of a friend who you met for lunch. At some point in the conversation they mention that they have a startup and are seeking someone like you. This revelation colors the whole conversation from that point onward. Indeed, each sentence colors the conversation from moment to moment. But, you reasonably respond, we can’t test that sort of C, modern AIs don’t do even ELIZA-level dialog yet! What’s the phrase??? “I rest my case?”
- darawk 4y agoIt's a lame test, but I don't think most people were claiming that it proves general compositionality. What it does prove is that compositionality is possible with these models, and will likely improve rapidly, as everything else has that they've gotten a toehold into. Ironically, the very fact that there is now a compositionality benchmark, as Gary points out, is all you really need to know that it's going to fall in the next decade, and probably sooner than that. I'm not aware of any major benchmark dataset upon which enormous progress has not been made in the last few years. And i'd be more than willing to bet anyone anything they'd like that a great deal of progress will be made on this one over the next few.
- abrax3141 4y agoMarcus seems to be treating it as the hallmark of intelligence (I think he actually uses that phrase), so arguing about whether the hack manages to get the tree into the effective object slot vs the effective subject slot is really not much of a hallmark.
- kcorbitt 4y agoWhat kind of insights do you expect a machine to be able to extract? I passed your example to GPT-3, and got back results that seem about the same as I'd expect from a human: PROMPT: > This is a test of reading comprehension. Read the following passage and answer the questions below in order. > Passage: > "You are engaged in a long conversation with someone, perhaps a friend of a friend who you met for lunch. At some point in the conversation they mention that they have a startup and are seeking someone like you. This revelation colors the whole conversation from that point onward. Indeed, each sentence colors the conversation from moment to moment" > Questions: > 1. What is the "revelation" referenced? > 2. What do you think the person is hoping to achieve by inviting you to lunch? > Answers: GENERATED OUTPUT: > 1. The revelation is that the person has a startup and is seeking someone like the reader. > 2. It is possible that the person is hoping to recruit the reader for their startup.
- daveguy 4y agoNow ask it a question.
- aaroninsf 4y agoSo many trees, so little forest. Gary Marcus comes off in this as very long on pious snark and very short on awareness of his own vulnerability to cognitive error, which is just as striking as any of his targets. The error in his question being: unconsidered linear extrapolation in a domain that is demonstrably non-linear, indeed exponention. To frame this a different way, he's very pious for maintaining a faith in his specific god ("strong AI is like production fusion power, ten to twenty year from now for every now"), but he's worshiping a god of the gaps. The gap in this case being <checks notes> "compositionality." Yes, language is hard. Yes, strong AI isn't here. But to not take a hard look at the jump up the abstraction hierarchy going on with contemporary ML and not nervously wonder if your faith is maybe a little too sure for a "scientist"...? Bad look when you're on the offensive.
- wrycoder 4y agoJust keep laughing. I'd like to hear Ray Kurzweil's view (he's working at Google and is awfully quiet.) Human consciousness is over-rated. I'm reminded of Minsky's Society of Mind - a number of separate, communicating systems. To me, that sounds a lot like what is going on in Google, but they are hiding that.
- dmix 4y agoRay Kurzweil was always taken as a bit of a loon and over-optimistic. Even way back at the peak of his popularity a decade ago. Just look at any of the old HN threads, anyone paying attention would have noticed. He's still a useful mind to have around. Like having scifi authors and philosophers. They don't have to be completely grounded in reality to provide useful projections as sources of inspiration and to challenge our grasp of history and growth.
- wrycoder 4y agoHired ten years ago at Google as Director of Engineering [0], His book, released that year, was "How to Create a Mind". He's still there, and that's what he's doing, I think. He is supposed to release his new book, "The Singularity is Nearer" in 2022, according to his website. I'll be reading that! [0] https://www.wsj.com/articles/BL-DGB-25711 https://www.wsj.com/articles/BL-DGB-25711
- wrycoder 4y agoOh, he just did an interview with Lex Friedman: https://youtu.be/ykY69lSpDdo https://youtu.be/ykY69lSpDdo
- stephc_int13 4y agoWe have absolutely no way to tell how far from "AGI" we are. What we know for sure is that we're not there yet. And what seems likely is that we're getting closer, and that's something. That is as much prediction we can get. I don't think that Compositionality is a wall, it is clearly an interesting feature, but I think that it is pretty clear by now that the Turing test or anything in the same spirit is far from sufficient.
- emmelaich 4y ago"a lightbulb surrounding some plants" is a weird phrase and a human feeling pedantic might well come up with the picture shown. A more typical phrase would be "lightbulbS around some plants" - note the plural. Maybe I'm missing something but using non-typical language won't work when it's been trained on normal language.
- ivanbakel 4y agoI think you've misunderstood that example in the article. The AI isn't being asked to generate an image from the prompt, it's being asked to match the similar prompts to the different images. Winoground is basically a reading-comprehension test suite, which links back to the point made in the article that AI can't handle non-typical language precisely because it lacks reading comprehension (or any semantic model of language.) As the article points out, human runs of Winoground manage to match the vast majority of prompts to the correct image, so it's not a question of atypical language being too hard to understand. You may want to also read the author's other article[0] about the lack of semantic comprehension in AI models. 0: https://garymarcus.substack.com/p/horse-rides-astronaut https://garymarcus.substack.com/p/horse-rides-astronaut
- dane-pgp 4y ago"A lightbulb. Surrounding: some plants." https://frinkiac.com/img/S07E18/562995.jpg https://frinkiac.com/img/S07E18/562995.jpg
- darawk 4y agoEvery concrete prediction Gary has made has been falsified. All of his others are insufficiently precise to be falsified. His GPT-2 examples were thoroughly defeated by GPT-3. Horse riding astronaut is solved. Neural knowledge graphs are a successful thing now. Compositionality isn't solved, but progress is clearly being made. If he was a serious person, this post could have been a few sentences: "No neural network will achieve <x> score on <y> metric on the Winoground dataset within the next <n> years". Simple, concrete, falsifiable. He has not done this, and one has to wonder why.
- DisjointedHunt 4y ago[flagged]
- SergeAx 4y ago> Full disclosure, I read Alexander’s successor Slate Star Codex, Astral Codex Ten, myself, and often enjoy it…when, that is, he is not covering artificial intelligence, about which we have had some rather public disagreements. Can it be a case of Gell-Mann Amnesia Effect? (https://en.m.wikipedia.org/wiki/Michael_Crichton#GellMannAmnesiaEffect https://en.m.wikipedia.org/wiki/Michael_Crichton#GellMannAmn...)
- water8 4y ago
- SilverBirch 4y agoI often hear on places like here that Scott Alexander is interesting and deep and insightful. But then I see bits and pieces like this. This blogger doesn't need to go into some deep analysis of compositionality to go "You came up with a 5 question test and decided 1 answer out of 10 attempts would be a pass". We've gone from 90%+s in imagenet to this as a pass mark? It's like sure we can dissect all the statistical risks of this, but why bother? It's self evident bullshit. You might as well have just posted a link to Scott Alexander's original blog claiming victory with just "Lol ok". Just post a screenshot of the phrase "An oil painting of a robot in a factory looking at a cat wearing a top hat", show the pictures of a robot near a cat that has a top hat, not in a warehouse, and say "lol ok."
- phreeza 4y agoI feel like he has kind of lost his spark a bit, but he does draw an interesting group of commenters. Similar to hacker news in that regard, sometimes the linked articles are a bit mundane but there is gold in the comments.
- ravi-delia 4y agoThis blog misses the point; he made a bet, and the people on the other side also accepted the terms. Nowhere did Alexander claim composition was a solved problem, just that the terms of the bet were satisfied. Generative models are still bad at composition, but claiming they will literally never improve requires some amount of additional evidence
- Viliam1234 4y agoExactly. The bet wasn't that AI will do composition correctly all the time, but that it will do composition correctly sometimes. That is how both sides understood it. The analogy with a student at exam misses the point. If you do art -- even as a human artist -- you do not need a 100% success rate. A 10% success rate is okay if you are willing to simply throw away the remaining 90% of the pictures. If you have an AI that at a click of a button can generate 10 beautiful pictures, 1 of them containing exactly what you wanted, that just means you need to make two clicks in order to get the picture you wanted. That is an awesome thing.
- concinds 4y ago"A lightbulb surrounding some plants" is not English. If a wolf pack is surrounding a camp, we understand what it means. If a wolf is surrounding my camp; does that mean I'm in his stomach? Absurd. "A lightbulb containing some plants," makes sense, not "surrounding". It's too small to surround anything, which humans (and apparently, current AI) understand. Paradoxically, only primitive language models would actually understand the inverted sentences; proper AIs should, like humans, be confused by them; since zero human talks like that. The only reason the Huggingface people (in their Winoground paper) got 90% of humans "getting the answer right" with these absurd prompts because of humans' ability to guess what is expected of them by an experimenter. Do it in daily life instead of a structured test, and see if these same people get it right. It's exactly as if I gave you the sequence, in an IQ-test context: "1 1 2 3" and asked you to give me the next number. You'd give the Fibonacci sequence, because you know I expect it; no matter that it's a stupid assumption to make because the full sequence might as well be "1 1 2 3 1 1 2 3 1 1 2 3", and you don't have enough information to know the real answer. Do we really want AIs that similarly "guess" an answer they know to be wrong, just because we expect it? Or (in number sequence example) AIs that don't understand basic induction/Goodman's Problem? I'd like to add that the author, who keeps referring to himself as a scientist, is in fact a psychology professor. In his Twitter bio, he states that he wrote one of the "Forbes 7 Must-Read Books in AI", which discredits him as a fraud since Forbes can be paid to publish absolutely whatever you ask them to (it's not disclosed as sponsored content, and they're quite cheap, trust me).
- adamsmith143 4y ago>"A lightbulb containing some plants," makes sense, not "surrounding". It's too small to surround anything, which humans (and apparently, current AI) understand. Paradoxically, only primitive language models would actually understand the inverted sentences; proper AIs should, like humans, be confused by them; since zero human talks like that. Not sure this is credible. Most if not all human adults are capable of understanding what young children just learning to speak mean most of the time, not only people with very low IQs. So why would this be any different? Presumably the smarter the AI the better it can understand poor grammar.
- 4y ago
- theptip 4y agoI thought Scott Alexander jumped the gun a bit by declaring victory in this case, just because the prompts used were not the original ones (robot vs. person due to content filters). But Marcus is way off base here and sounding petulant; Alexander is clearly not claiming AI has solved compositionality, his claim is the much narrower one that he won his bet. And the general context to the bet is that usually when he writes an article on AI (at least for the last few years), someone says “we will never get X in the next 5 years”, Alexander makes a bet that it will happen sooner, and X always happens sooner. In this case the X was some loose low bar for the next iteration of compositionality above DALL-E 2 with a multi-year timeframe, and SOTA models at the time of the discussion could (arguably at least) meet that bar. Alexander’s broad claim on compositionality is that simply throwing more scale and/or data at the problem seems likely to solve the problem, to which Marcus counters that these models lack something fundamental and can’t be scaled to human performance. FWIW I find Marcus’ position to be a bit frustratingly ambiguous; he seems to blend two distinct positions: A) NN models are not a model for human intelligence/language B) NN models cannot reach AGI He seems to fluidly switch between these critiques in a way I find a bit irritating. I think it’s quite clear that NN architectures have little to do with the way the human brain does language understanding, lacking the gross structure of the brain, which is certain to affect cognitive capabilities and tendencies. So A) is trivially true. But no AI maximalist cares about using these models as a way to understand or model human language. They care about general intelligence. Even granting A), that does nothing to prove B). Perhaps he simply believes B requires A? That would be odd but would explain his approach.