8 ms·
Drawing a childrens story with DALL-E 2
- TheMagicHorsey 4y agoI wish there was a program that allowed you to draw some input pictures/shapes, and the program takes that input and creates a picture in a given style. That would be a gamechanger for storytelling.
- yojo 4y agoNVIDIA did something like that for landscapes: https://www.nvidia.com/en-us/studio/canvas/ https://www.nvidia.com/en-us/studio/canvas/
- YeGoblynQueenne 4y agoThat's a really great idea! If we train our kids from a young enough age to not notice the incoherence in content generated by systems like DALL-E they will probably never learn to notice it. Then, when they're adults, they won't mind it! And we can all pretend it's not even there! (e.g. why does the little girl keep changing hair, face, clothes, age, ....)
- ffhhj 4y agoThe first rule of the utopia is not questioning the utopia ;)
- deleted 4y ago[deleted]
- renewiltord 4y agoOr perhaps they'll learn the concept / instance separation earlier because they will comprehend the notion of an image depicting something without needing continuity of characters to express it.
- YeGoblynQueenne 4y agoOh, I'm sure it's possible to understand something and still be annoyed by it. For example, an incoherent story. I think that will piss off most people.
- glutamate 4y agoHow do you ensure a continuity of the main characters with a DALL-E generated comic? The example here doesn't seem to do this well. There characters are different in each frame
- capableweb 4y agoI don't have access so not sure it would work, but maybe the author could be more specific with the prompt to get a character that looks more similar for each frame. "Girl with brown hair, small nose, blue eyes, white shirt and blue skirt goes inside school" for example.
- nutanc 4y agoI tried this. Every time a new girl pops up :)
- capableweb 4y agoHm, that's sad. Maybe in the future you'll be able to assign specific results into variables (like saving one result of this avatar you generated as $SCARED_OF_SCHOOL_GIRL) and enforce that avatar to be used in the future, would help in these cases.
- bredren 4y agoIf it was run enough times you would get similar enough looking girls which could be used! Just need to train it on that girl somehow? This is the infinite universes thing right?
- jonahx 4y agoThe pictures, taken one by one, are impressive. Is it possible to tell DALL-E to use the same girl in each panel? That is the only detail preventing me from being convinced this isn't a real children's book.
- monkeydust 4y agoJust tried, I was curious. Answer is yes but not very well. So you can generate a picture say 'picture of a boy playing with a ball' then go into edit mode (think this is new) and erase ball, then change prompt to 'picture of a boy playing with toy car'. It will keep the original elements of boy not erased and put in a toy car. The character is kept the same but doesn't work as well as you might think as the body, face pose are the same. Still, you can see this getting better over time. Only got access this morning, hard not be impressed.
- nutanc 4y agoI tried a lot, but thats one shortcoming. You can try editing and redrawing. Some success is there, but scope for improvement is there.
- bredren 4y agoSeems like this would seriously enhance this product. As a comparison, stock photos and video of people are often available as a scene showing the same folks in a bunch of scenarios and angles. You kind of need this to build something of any complexity out of the art.
- GistNoesis 4y agoI don't know if it's possible yet, but the technology behind (guided diffusion), allows it. But you will probably have to train it with a dataset that shows multiple associated images together. Denoising diffusions have been successfully applied to produce coherent videos. In fact, there are even some benefits to train on video because images share information between them, which make context extraction easier. Counterintuitively, because of the shared information between frames, their computations can also be partially shared, it can be almost free to generate multiple frames at the same time : A pair of RGB image can be considered as a 6D image, a 5-frame sequence a 15D image, and the neural network can learn to compress the various image channel together. The subsequent neural network layers can be kept 256D (for example) the same size as they were for a single frame. Of course this trick won't work if the pair of image are too different but you can always stack images side by side to have a bigger film-roll image and you use a standard attention mechanism to attend to various parts.
- ffhhj 4y agoSeems useful to create quick sketches for an artist to re-draw the scenes so they keep a given style. I can create 3D models from sketches, but don't have much creativity to compose the scenes, a tool like DALL-E would give me a big kick-start.
- rebuilder 4y agoI’m not entirely sure about that. It feels like saying you could just use an AI tool to sketch out a piece of music and then have a composer fix it up. I suspect it will turn out to be much more work for the artist to find what it is you actually liked about the generated output and rework the pictures based on that, rather than the usual process where the client describes what they need and the artist uses their experience and understanding of the human mind to decide how to represent that. I guess I’m saying that the art in “art” is not the superficial skill of making pretty pictures, but the concurrently honed skill of making meaningful choices. I’m not sure you can just bypass learning one without compromising the other.
- manimino 4y agoIt ought to be possible to guide a diffusion-based model to make several panels about a single, consistent, AI-invented character. It can already generate several independent images containing known characters that are in the training data (Obama, Pikachu, etc.)
- jcims 4y agoI'm a total layperson in this area but there used to be a concept of 'fine tuning' with language models (may still be). My understanding of those is that you could inject additional training data that overlays(?) the existing model in some way to help direct the output. In this case it seems that you could fine-tune the generative image model with previous frames to provide that continuity. After all that's what we do when we read the panel, we instantly store the previous one in memory so that we can actually recognize the difference in the next panel.
- mistrial9 4y agoas a trained artist - I find this cancerous. Is it not obvious that some extreme computer users are explicitly going into the most sensitive and therefore fragile aspects of human life.. like a thrill seeker addict going for more and more intensity.. major no
- deleted 4y ago[deleted]
- ALittleLight 4y agoI'm sure professional drivers don't like the idea of self-driving cars either.
- mistrial9 4y agono - not the same .. what you say is machinery utilitarian, for pay
- ALittleLight 4y agoI am neither a professional driver nor a professional artist but I have paid both for their services. In both cases it's an exchange, money to take something somewhere or money for the illustrations I want. They both, from my perspective as a user, could be replaced by software that performs well. I'm reading your comment as saying that there is some sine qua non about art that professional driving lacks - but I think that may just be your bias as an artist. If software produces results that are as good or better than yours, what is lost by replacing artists with software? I'm also not writing this just to dig at professional artists and drivers. I think my career too is in the process of being replaced by software. I have doubts and misgivings about whether this will be entirely good, but it seems clear that it is happening and that most (all?) professions are or will soon be in a similar state.
- mistrial9 4y ago>what is lost by replacing artists with software ? computers are useless, they can only provide answers
- guerrilla 4y agoUmmm... looks as bad as you'd expect? What is there to talk about?
- tartoran 4y agoUmm.. another somewhat lucrative profession killed by tech?
- djmips 4y agoYou're hedging about it being lucrative. I can tell you it's not. Also this is pretty terrible. It'll need a lot of improvement.
- tartoran 4y agoOk, even if it wasn’t very lucrative it most likely to become even less so.
- fullshark 4y agoWe don't say that here, we say it's another somewhat lucrative profession "democratized."
- djmips 4y agoI'm glad it's not just me that looks at this kind of thing and thinks it's awful.
- baq 4y agoyeah exactly, I wouldn't want to read this book to my children. it is an alpha version of the future hence good content for hacker news, not necessarily something that should go into a real children's book ;)
- tyingq 4y agoHuman captions on the generated pictures are much better. https://pbs.twimg.com/media/FVomoVbXoAI5pCB?format=jpg https://pbs.twimg.com/media/FVomoVbXoAI5pCB?format=jpg
- awillen 4y agoNaturally the first comments here are criticisms that the characters change from frame to frame, but I think that's worth putting aside, since any rational person will understand that systems like DALL-E will have the ability to maintain a level of continuity between images in the near future (if it's not already there and just not exposed). Besides that, this is pretty good - certainly a few things that don't seem quite ideal (the fifth frame doesn't seem to capture the meaning of the text especially well), but enough to make you think that in not too long, it is fully plausible that DALL-E will be fully able to act as an illustrator for this use case. On the one hand, pretty exciting. On the other, certainly harrowing for children's book illustrators. I wonder how long it will be before I can use this for my marketing emails. I sell dog treats, and I have a fairly simple template with an image at the top. That's almost always a photo of my products and/or my dogs. How long before I can just ask for an image from DALL-E for something like "Dog sitting next to grill in back yard with American flags and other patriotic decorations" for the top of my Fourth of July email? How long before I can feed it a picture of my products and have it generate photorealistic images of dogs eating them, thus replacing the photographer I use for product shoots? It's an exciting prospect for me as a small business owner - lots of time and money saved in an area where I'm not an expert - but definitely pretty scary for people who create visual imagery of any kind.
- nutanc 4y agoDalle2 is really good for stock marketing content. Your use case is certainly feasible. Especially with the edit feature we can make dogs eat your product. If you can DM me a product image I can show a POC :)
- synu 4y agoMaybe you can set it aside as a tech demo, but the ask to please show this to your children feels somehow creepy to me. They are still so young and impressionable, and providing them with perceptibly incoherent stories and imagery just seems.. off.
- capableweb 4y ago> providing them with perceptibly incoherent stories and imagery just seems.. off Not sure it really matters, half of all children books are filled of incoherent stories and imagery, but I think it's mostly adults who notice that.
- ducktective 4y agoI suggested a similar thing and couple of people dropped in with more specialized knowledge on this: https://news.ycombinator.com/item?id=31425690 https://news.ycombinator.com/item?id=31425690
- jcims 4y agoHey nutanc did you create the text or is that generated as well? I had to mess around with the prompt a little but GPT-3 just did this: Prompt: Hi i'm Elsie, a first grader going to school for the first time. This is a six panel comic speaking in the first person about how my day went. I'll include a suggestion for the illustrator to draw each frame in square brackets: Generated text: [Elsie is standing in front of her house with her backpack on, looking a little nervous.] Hi, I'm Elsie. Today is my first day of school. I'm a little nervous, but I'm excited too. [Elsie is walking to school with her mother.] I walk to school with my mom. We talk about what my day might be like. [Elsie is standing in front of her school with her mother, who is giving her a hug and kiss goodbye.] My mom gives me a hug and a kiss, and then I'm off! [Elsie is in her classroom, sitting at her desk.] I sit at my desk and the teacher starts talking. I learn a lot of new things today. [Elsie is standing in front of her house again, this time with a big smile on her face.] I had a great first day of school! I can't wait to go back tomorrow. Any time i mess with this stuff i get lost trying different variations. Here's a few more (# marks the prompt) https://pastebin.com/deC0RxDb https://pastebin.com/deC0RxDb
- renewiltord 4y agoRemarkable. One could imagine the Primer from The Diamond Age actually existing in the near future.
- nutanc 4y agoThe text is picked up from https://storyweaver.org.in/stories/98-rani-s-first-day-at-school https://storyweaver.org.in/stories/98-rani-s-first-day-at-sc... for this example. But I have experimented with GPT3 generating content and Dalle generating the image. For example here, https://twitter.com/nutanc/status/1538372995764617216 https://twitter.com/nutanc/status/1538372995764617216
- deleted 4y ago[deleted]
- blacksmith_tb 4y ago
- random_upvoter 4y agoThis is very impressive, technically. Surely looks like AI is going to be the death of mediocre artists.
- ALittleLight 4y agoI was just fantasizing about how to do stuff like this. In my fantasy there were two DALL-E-like models. One for generating characters and one for generating scenes and a system for adding a generated character to a scene. The user describes a character to the character model and gets candidate illustrations. The user can then save a character by giving it a name and use the name to generate variations of the same character. I'm imagining that either you just have a lot of sliders to vary the vector that the character is drawn from or maybe it's possible to combine ideas (e.g. Character X + running). Then the user just illustrates the scenes they want, the characters, combines them together, and gets a useful image.
- kazinator 4y agoOh goodie, here comes the deluge of low-effort children's books.
- capableweb 4y agoYeah, because that would definitely be new! There are already a TON of low-effort children's book, I'm not sure if DALL-E generated ones would be a improvement or not, I've definitely seen many that are worse than this.
- TheMagicHorsey 4y agoI wish there was a program that allowed you to draw some input pictures/shapes, and the program takes that input and creates a picture in a given style. That would be a gamechanger for storytelling.
- m3kw9 4y agoWhile the drawing looks decent, but not impressed with the story telling part of the pictures. Didn’t give that emotional feel