3 ms·
Yup it's an arbitrary metric, but I tried cajoling various image models into generating spork pictures with highly detailed descriptions (I have ComfyUI & AUTOM
by quesomaster9000 2y ago
Yup it's an arbitrary metric, but I tried cajoling various image models into generating spork pictures with highly detailed descriptions (I have ComfyUI & AUTOMATIC1111 locally, and many models), which lead to me creating the site.
I'd say a better test for adherence is how well a model does when the detailed description falls in between two very well known concepts - it's kinda like those pictures from the 1500s of exotic animals seen by explorers drawn by people using only the field notes after a long voyage back.
- vunderba 2y agoHere's a test against the full Flux Dev model where we try to describe the physical appearance of a spork without referencing it by name. Prompt: A hybrid kitchen utensil that looks like a spoon with small tine prongs of a fork made of metal, realistic photo https://gondolaprime.pw/pictures/flux-spork-like.jpg https://gondolaprime.pw/pictures/flux-spork-like.jpg The combination of T5 / clip coupled with a much larger model means there's less need to rely on custom LoRAs for unfamiliar concepts which is awesome. EDIT: If you've got the GPU for it, I'd recommend downloading a copy of the latest version of the SD-WEBUI Forge repo along with the DEV checkpoint of Flux (not schnell). It's super impressive and I get an iteration speed of roughly 15 seconds per 1024x1024 image.
- vunderba 2y agoWelll.... there's a hundred ways we could measure prompt adherence everything from: - Descriptive -> describing a difficult concept that is most certainly NOT in the training data - Hybrids -> fusions of familiar concepts - Platonic overrides -> this is my phrase for attempting to see how well you can OVERRIDE very emphasized training data. For example, a zebra with horizontal stripes. etc. etc.