3 ms·
I'm not sure why one would even strike metal against a crucible! It's a container for liquid metal. One of the outputs shows it being smashed by the manoeuvre,
by speerer 6d ago
I'm not sure why one would even strike metal against a crucible! It's a container for liquid metal. One of the outputs shows it being smashed by the manoeuvre, which is probably the most realistic outcome of all of them.
Sorry, I'm not trying to nitpick. I'm just joining in because I'm interested in how the models dealt with the request.
- vunderba 6d agoWell this is HN - original home of the "ummm actually..." - so I appreciate when people pick all the nits. :) Even though I prompted for a crucible in the prompt, I think the fact that the prompt also contained terms like “blacksmith” and “hammer,” caused it to lean towards anvils over crucibles in some of the pictures (which as you brought up makes more sense anyway).
- seemaze 5d agopedants unite!
- apothegm 5d agoPerhaps for the particular image you liked. But choosing something “typical” or average over what it was instructed to do is a massively common failure mode for AI. One that makes the difference between a useful model and one that makes you want to throw your laptop out a window.
- vunderba 5d agoI agree. In fact the entire reason I initially built GenAI Showdown was because many of the comparative tests on places like Image Arena Leaderboard [1] are not designed to challenge models on prompt adherence. Even when they are, a considerable number of amateur judges tend to prioritize aesthetics over adherence or instruction-following. I'll likely be redoing that particular bench with added minimum passing criteria of an anvil. [1] - https://huggingface.co/spaces/ArtificialAnalysis/Text-to-Image-Leaderboard https://huggingface.co/spaces/ArtificialAnalysis/Text-to-Ima...
- RugnirViking 5d agoive been taking a look there. I think for some of the image to image ones you really ought to use real images as the base prompt, there are a lot of weird things happening where the base image has issues and then its hard to say if the model should be correcting the flaws or not. Things like "childrens drawing app where each crayon is clearly sized to be tapped on" but the base image has the rainbow cut off partway across with only half the red and black crayons visible. Also "vintage" photography being corrected, where the original is clearly ai generated with unrealistic sharp focus everywhere