8 ms·
I don't see how this is "embarrassing" in the slightest. These models are not human brains, and the fact that people equate them with human brains is an embarra
by tensor 2y ago
I don't see how this is "embarrassing" in the slightest. These models are not human brains, and the fact that people equate them with human brains is an embarrassing failure of the humans more than anything about the models.
It's entirely unsurprising that there are numerous cases that these models can't handle that are "obvious to humans." Machine learning has had this property since its invention and it's a classic mistake humans make dealing with these systems.
Humans assume that because a machine learning model has above human accuracy on task X that it implies that it must also have that ability at all the other tasks. While a human with amazing ability at X would indeed have amazing abilities at other tasks, this is not true of machine learning models
The opposite thinking is also wrong, that because the model can't do well on task Y it must be unreliable and it's ability on task X is somehow an illusion and not to be trusted.
- cs702 2y agoIt is embarrassingly, shockingly bad, because these models are advertised and sold as being capable of understanding images. Evidently, all these models still fall short.
- kristjansson 2y agoIt's surprising because these models are pretty ok at some vision tasks. The existence of a clear failure mode is interesting and informative, not embarrassing.
- knowaveragejoe 2y agoNot only are they capable of understanding images(the kind people might actually feed into such a system - photographs), but they're pretty good at it. A modern robot would struggle to fold socks and put them in a drawer, but they're great at making cars.
- pixl97 2y agoI mean, with some of the recent demos, robots have got a lot better at folding stuff and putting it up. Not saying it's anywhere close to human level, but it has taken a pretty massive leap from being a joke just a few years ago.
- startupsfail 2y agoHumans are also shockingly bad on these tasks. And guess where the labeling was coming from…
- simonw 2y agoI see this complaint about LLMs all the time - that they're advertised as being infallible but fail the moment you give them a simple logic puzzle or ask for a citation. And yet... every interface to every LLM has a "ChatGPT can make mistakes. Check important info." style disclaimer. The hype around this stuff may be deafening, but it's often not entirely the direct fault of the model vendors themselves, who even put out lengthy papers describing their many flaws.
- jazzyjackson 2y agoThere's evidently a large gap between what researchers publish, the disclaimers a vendor makes, and what gets broadcast on CNBC, no surprise there.
- jampekka 2y agoA bit like how Tesla Full Self-Driving is not to be used as self-driving. Or any other small print. Or ads in general. Lying by deliberately giving the wrong impression.
- verdverm 2y agoIt would have to be called ChatAGI to be like TeslaFSD, where the company named it something it is most definitely not
- deleted 2y ago[deleted]
- fennecbutt 2y agoWhy do people expect these models, designed to be humanlike in their training, to be 100% perfect? Humans fuck up all the time.
- TeMPOraL 2y agoThey're hardly being advertised or sold on that premise. They advertise and sell themselves, because people try them out and find out they work, and tell their friends and/or audiences. ChatGPT is probably the single biggest bona-fide organic marketing success story in recorded history.
- foldr 2y agoThis is fantastic news for software engineers. Turns out that all those execs who've decided to incorporate AI into their product strategy have already tried it out and ensured that it will actually work.
- ben_w 2y ago> Turns out that all those execs who've decided to incorporate AI into their product strategy have already tried it out and ensured that it will actually work. The 2-4-6 game comes to mind. They may well have verified the AI will work, but it's hard to learn the skill of thinking about how to falsify a belief.
- TeMPOraL 2y agoYou mean this one here? - https://mathforlove.com/lesson/2-4-6-puzzle/ https://mathforlove.com/lesson/2-4-6-puzzle/ Looking at the example patterns given: MATCH 2, 4, 6 8, 10, 12 12, 14, 16 20, 40, 60 NOT MATCH 10, 8, 6 If the answer is "numbers in ascending order", then this is a perfect illustration of synthetic vs. realistic examples. The numbers indeed fit that rule, so in theory, everything is fine. In practice, you'd be an ass to give such examples on a test, because they strongly hint the rule is more complex. Real data from a real process is almost never misleading in this way[0]. In fact, if you sampled such sequences from a real process, you'd be better off assuming the rule is "2k, 2(k+1), 2(k+2)", and treating the last example as some weird outlier. Might sound like pointless nitpicking, but I think it's something to keep in mind wrt. generative AI models, because the way they're trained makes them biased towards reality and away from synthetic examples. -- [0] - It could be if you have very, very bad luck with sampling. Like winning a lottery, except the prize sucks.
- mrbungie 2y agoThese models are marketed as being able to guide the blind or tutoring children using direct camera access. Promoting those use cases and models failing in these ways is irresponsible. So, yeah, maybe the models are not embarrasing but the hype definitely is.
- cs702 2y ago> Promoting those use cases and models failing in these ways is irresponsible. Yes, exactly.
- deleted 2y ago[deleted]
- scotty79 2y agoYou'd expect them to be trained on simple geometry since you can create arbitrarily large synthetic training set for that.
- sfink 2y agoWell said. It doesn't matter how they are marketed or described or held up to some standard generated by wishful thinking. And it especially doesn't matter what it would mean if a human were to make the same error. It matters what they are, what they're doing, and how they're doing it. Feel free to be embarrassed if you are claiming they can do what they can't and are maybe even selling them on that basis. But there's nothing embarrassing about their current set of capabilities. They are very good at what they are very good at. Expecting those capabilities to generalize as they would if they were human is like getting embarrassed that your screwdriver can't pound in a nail, when it is ever so good at driving in screws.
- insane_dreamer 2y ago> is an embarrassing failure of the humans more than anything about the models No, it's a failure of the companies who are advertising them as capable of doing something which they are not (assisting people with low vision)
- simonw 2y agoBut they CAN assist people with low vision. I've talked to someone who's been using a product based on GPT-4o and absolutely loves it. Low vision users understand the limitations of accessibility technology better than anyone else. They will VERY quickly figure out what this tech can be used for effectively and what it can't.