4 ms·
Perhaps for the particular image you liked. But choosing something “typical” or average over what it was instructed to do is a massively common failure mode for
by apothegm 5d ago
Perhaps for the particular image you liked. But choosing something “typical” or average over what it was instructed to do is a massively common failure mode for AI. One that makes the difference between a useful model and one that makes you want to throw your laptop out a window.
- vunderba 5d agoI agree. In fact the entire reason I initially built GenAI Showdown was because many of the comparative tests on places like Image Arena Leaderboard [1] are not designed to challenge models on prompt adherence. Even when they are, a considerable number of amateur judges tend to prioritize aesthetics over adherence or instruction-following. I'll likely be redoing that particular bench with added minimum passing criteria of an anvil. [1] - https://huggingface.co/spaces/ArtificialAnalysis/Text-to-Image-Leaderboard https://huggingface.co/spaces/ArtificialAnalysis/Text-to-Ima...
- RugnirViking 5d agoive been taking a look there. I think for some of the image to image ones you really ought to use real images as the base prompt, there are a lot of weird things happening where the base image has issues and then its hard to say if the model should be correcting the flaws or not. Things like "childrens drawing app where each crayon is clearly sized to be tapped on" but the base image has the rainbow cut off partway across with only half the red and black crayons visible. Also "vintage" photography being corrected, where the original is clearly ai generated with unrealistic sharp focus everywhere