5 ms·
Dumb question, but how are DALL-E's (and any other AI generative algorithm) result's so.. smooth? For example, I could write a heuristic algorithm to product t
by sydthrowaway 4y ago
Dumb question, but how are DALL-E's (and any other AI generative algorithm) result's so.. smooth?
For example, I could write a heuristic algorithm to product the same thing using a Google image search, but it would look like MS word clip art.
- Enginerrrd 4y agoLots of denoising steps after the initial attempt at forming a connection to the prompt is made?
- sillysaurusx 4y agoThis is one of my favorite topics in all of AI. It was the most surprising and mysterious discovery for me. The answer is that the training process literally has to make the results smooth. That’s how training works. Imagine you have 100 photos. Your job is to classify them by color. You can place them however you want, but similar colors should be physically closer together. You can imagine the result would look a lot like a photoshop RGB picker, which is smooth. The surprise is, this works for any kind of input. Even text paired with images. The key is the loss function (a horrible name). In the color picker example, the loss function would be how similar two colors are. In the text to image example, it’s how dissimilar the input examples are from each other (Contrastive Loss). The brilliance of that is, pushing dissimilar pairs apart is the same thing as pulling similar pairs together, when you train for a long time on millions of examples. Electrons are all trying to push each other apart, but your body is still smooth. The reason it’s brilliant is because it’s far easier to measure dissimilar pairs than to come up with a good way of judging “does this text describe this image?” — you definitely know that it isn’t a bicycle, but you might not know whether a car is a corvette or a Tesla. But both the corvette and the Tesla will be pushed away from text that says it’s a bicycle, and toward text that says it’s a car. That means for a well-trained model, the input by definition is smooth with respect to the output, the same way that a small change in {latitude,longitude} in real life has a small change in the cultural difference of a given region of the world.
- Michelangelo11 4y agoDo you by any chance have a link to a paper or article that explains this in detail? I'd love to understand it better.
- jeabays 4y ago
- sillysaurusx 4y agoIt doesn’t exist. The above explanation is the result of me spending almost all of my time immersing myself in ML for the last three years. gwern helped too. He has an intuition for ML that I’m still jealous of. Your best bet is to just start building things and worry about explanations later. It’s not far from the truth to say that even the most detailed explanation is still a longform way of saying “we don’t really know.” Some people get upset and refuse to believe that fundamental truth, but I’ve always been along for the ride more than the destination. It’s never been easier to dive in. I’ve always wanted to write detailed guides on how to start, and how to navigate the AI space, but somehow I wound up writing an ML fanfic instead: https://blog.gpt4.org/jaxtpu https://blog.gpt4.org/jaxtpu (Fun fact: my blog runs on a TPU.) I’m increasingly of the belief that all you need is a strong desire to create things, and some resources to play with. If you have both of those, it’s just a matter of time — especially putting in the time. That link explains how to get the resources. But I can’t help with how to get a desire to create things with ML. Mine was just a fascination with how strange computers can be when you wire them up with a small dose of calculus that I didn’t bother trying to understand until two years after I started. (If you mean contrastive loss specifically, https://openai.com/blog/clip/ https://openai.com/blog/clip/ is decent. But it’s just a droplet in the pond of all the wonderful things there are to learn about ML.)
- Michelangelo11 4y agoThanks! Really appreciate the response.
- sizzle 4y ago