3 ms·
All I see is mention of how various models generate image of "pelican riding bicycle(s)"
by pineapple_opus 5mo ago
All I see is mention of how various models generate image of "pelican riding bicycle(s)"
- emil-lp 5mo agoYes, the "pelican riding a bicycle" is the ultimate test of not understanding how LLMs work. Well, a combination of that and believing that replication of test data is a good measure of progress.
- vessenes 5mo agoSpicy — why does it show ultimate non-understanding?
- JohnKemeny 5mo agobecause success comes from reproducing a memorized pattern rather than transferable reasoning? At the same time failure proves little because most humans also could not manually create a correct SVG of a pelican riding a bicycle. What is it exactly that such a test is testing? In which situation would you measure the "competence" of a human being by asking them to write an SVG of a pelican riding a bicycle?
- okamiueru 5mo ago> most humans also could not manually create a correct SVG of a pelican riding a bicycle. Most humans absolutely can write this with a suitable vector graphics tool such as inkscape or illustrator. Surely, you're not suggesting that a fair comparison would be using a text editor? If so, would you suggest an equivalent raster based task would only be fair, if the human would manually assigning RGB values to each pixel?
- ClikeX 5mo agoWe all know the true test of AI is Will Smith eating spaghetti.
- ActionHank 5mo agoWait, are you saying you don't handcraft svgs of pelicans riding bicycles?