3 ms·
> I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican
by Wowfunhappy 2mo ago
> I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle".
Oh come now. I am extremely confident that if I hired a professional artist to draw a picture of a pelican riding a bicycle, I would get something inarguably much better than what today's best coding LLMs can produce.
- skygazer 2mo agoAt that rate you could use an image model, which was designed for the task. Thats the absurdity of this test. It’s often a text only generation model that has never seen a pelican, coerced into creating an xml graphical representation that a human might recognize. That it does anything passable is already astounding.
- Wowfunhappy 2mo agoI agree, it's hard! That's why it's still a good benchmark.
- techpression 2mo agoIt probably has a few million inputs on how a pelican and bicycle looks, not to mention the amount of data on how to create SVG’s. Ask it to create a relaxing spa website and it will, even though it has never seen a spa.
- didibus 2mo agoMost of the models are multi-modal and trained on images no? That's what they claim at least.
- skygazer 2mo agoYou’re right. Modern frontier models are now multimodal. I used often as weak a hedge, because I know at least his gpt3.5 turbo and llama3.1 generated pelicans were from text only models without image training. The chinese models are interesting, because before their vision models existed they may have been distilling text only models from text output of American vision models, so they could have benefited from the teacher model’s vision capability without being vision models themselves.
- pj_mukh 2mo ago> if I hired a professional artist to draw a picture of a pelican riding a bicycle, I think AI folks have done a terrible job of communicating this, but replacing a professional simply isn't the point. The point is to serve all the situations where people would've never considered hiring a professional, and where perfection or artistic merit isn't the point (say a personal throwaway recreation of an LOTR world). And I think in that regard the benchmarks are pretty good.
- Wowfunhappy 2mo ago> replacing a professional simply isn't the point. I'm not saying it is, just that there's obviously still room for the models to improve on this task.
- YmiYugy 2mo agoIMHO the output is bad enough that I can't imagine a use case for illustrations of this kind.