8 ms·
Is this just a new AI ghost story, or is there actually something scientifically repeatable here? Can we please get real nerdy about this? The article broadly
by schaefer 4y ago
Is this just a new AI ghost story, or is there actually something scientifically repeatable here? Can we please get real nerdy about this?
The article broadly describes a series of prompts, but do we have enough information figure out which AI engine was used, reverse engineer some likely prompts, and try to produce similar results (not exactly the same, as that may not be possible with AI prompts)?
Is it even possible to ask an AI image generator to "produce the opposite" of a prompt?
Is this just an RTFM moment (for me)? or is "producing the opposite" a misunderstanding of how weights work. I only have experience in midjourney, but my understanding is that with midjouney you can weight various prompts as ratios. For example I can build up a prompt to generate an image that is 2 parts "autumn landscape" and 3 parts "birthday cake". But with ratios, isn’t it true that negative weights just discarding that prompt? They don’t produce “the opposite”, right?
- schaefer 4y agofrom Supercomposite's original twitter post, we can see they claim to start off with a negatively weighted prompt of "Brando::-1" To me, this looks like midjourney prompts. At least we can say, that is valid midjourney syntax for sure, and it is using weights. but probably not to the effect the author portrays. for example, if I prompt "/imagine autumn landscape::2 birthday cake::3", I get this image[1]. But now if I tweak the prompt to "/imagine autumn landscape::-0.5 birthday cake::3", I get this image [2]. Critically, there is no trace of "opposite of an autumn landscape" in this image. It's all birthday cake... 100%. Indeed, this lines up with some midjouney documentation[3]. A negative weight will try to remove the thing in the prompt. BUT: what happens when we only have one prompt, and we use a weight to negate it in midjourney. i.e: "/imagine Brando::-1". At least for me, it get this error: "Invalid parameter The sum of all of the prompt weights must be positive" So I'm inclined to conclude that Supercomposite's post is more an act of creative story telling, than an accurate portrayal of their interactions with Midjourney. But I am still left wondering if there is a true "opposite of" operator in any AI image generator. [1]: https://cdn.discordapp.com/attachments/1006400576067739749/1017438050646761562/slowcamp_autumn_landscape_0253e597-b56e-4469-89b9-38ffe63a0675.png https://cdn.discordapp.com/attachments/1006400576067739749/1... [2]: https://cdn.discordapp.com/attachments/1006400576067739749/1017440309816340521/slowcamp_autumn_landscape_bd2939cd-a393-412f-870a-12b1e94dd851.png https://cdn.discordapp.com/attachments/1006400576067739749/1... [3]: https://midjourney.gitbook.io/docs/imagine-parameters#prompt-modifiers https://midjourney.gitbook.io/docs/imagine-parameters#prompt...
- yccs27 4y agoI mean, they might have modified the midjourney code to patch out the error message. Presumably the programmers put the check in because results for negative weight sums become nonsensical, not because it's impossible. Thinking in terms of classifiers, an image is almost never categorized as exactly 0% something, instead it's a positive value. Negative weights would make the net optimize for the smallest possible percentage. For the birthday cake, the negative weight is not strong enough to favor any "anti-autumn-landscape" patterns, only to remove features associated with autumn landscapes. But for weights that are all negative, it's plausible that the system will produce all the features that, in the training data, are anticorrelated with the prompt.
- rhn_mk1 4y agoThere's a comment thread diving into why the concept is so persistent and why it appears on negative queries: https://twitter.com/mattskala/status/1567300206969982979?t=CtaIv3LcTh9Wd_WyZ6_YXw&s=19 https://twitter.com/mattskala/status/1567300206969982979?t=C...
- TigeriusKirk 4y agoFollowing some stuff from that person, I came across this, a post trying to replicate Loab with various prompts. Shows visually what might be going on. https://twitter.com/chrysopoetics/status/1567673870433546242 https://twitter.com/chrysopoetics/status/1567673870433546242
- godelski 4y ago> Can we please get real nerdy about this? So as someone who works in generative modeling I'll give my best guess as to what is done and what is happening. It is a guess because they don't say everything, but there are some hints. Scambier linked these two twitter threads[0][1] which can give us some insight. > I'll explain negative prompt weights, in case you don't know. With these, instead of creating an image of the text prompt, the AI tries to make the image look as different from the prompt as possible. What's important here is that the machine doesn't actually know what the opposite is. In fact, I would ague we don't either. What you can do is use Lp distances from a latent representation. This is where things start to make sense. If in that we find faces as a large distance away, it is also unsurprising that we find many different facial characteristics. These first images look like there is a high mixture between strong masculine features and strong feminine features. These are not things we typically see in reality and combine with our hyperactive brains for recognizing other human faces, we enter the uncanny valley. Next I don't know if this was done on purpose or not, but there are very clear issues with scene lighting. I can totally believe that this is not on purpose because this is something generators are bad at already. So we have shadows cast along the face in unnatural ways. Upping the creepiness factors. Now we need to look at important features for recognizing faces: eyes, mouth, and nose. You may have noticed that text to image generators are typically really bad at these. Generators are also typically bad at facial symmetry (why we're trying to get transformers in, but this still isn't working to the degree we would like). In fact, I actually find it more interesting that these are coherent given the explanation of how the latent representation was created. So I think we have good explanations as to why this would turn creepy very fast. Especially given the hype and that the creator is leaning into it. But these are my best guesses. I can't really know without seeing what is done. But honestly, I am super interested and would like to see these latent representations and play around with them. This could be a good thing to investigate if you are trying to determine how smooth the latent manifold is, which is extremely important if we're going to make deeper content contributions and rely less on our prompt engineering. Maybe I'll have to play with some negative prompts (if I can find the time lol). [0] https://twitter.com/supercomposite/status/1567162288087470081 https://twitter.com/supercomposite/status/156716228808747008... [1]https://twitter.com/sheslostheplot/status/1567370919487893504 https://twitter.com/sheslostheplot/status/156737091948789350...
- cyanydeez 4y agoThis is all ghost stories. The internet is a schizophrenic and is deeply into examining itself