6 ms·
Using “underdrawings” for accurate text and numbers
- samcollins 5mo agoI found a simple technique to get reliable text and numbers in AI generated images. I’m surprised the image models aren’t already doing this, so wanted to share since I’m finding this so useful
- samcollins 5mo agoTLDR: use SVG to outline image correctly first, then send that image with your text prompt to get Gemini 3.0 Pro to render with correct numbers and text
- choppaface 5mo agoIsn’t this sort of just “chain of thought” (i.e. the seminal https://arxiv.org/abs/2201.11903 https://arxiv.org/abs/2201.11903 ) where the user is helping the model 1-shot or k-shot the solution instead of 0-shot? I’ve used a similar technique to great effect. I feel things are so new / moving so fast that it’s hard to have common lingo. So very helpful to have a blog / example! But I wonder if the phenomena has been seen / understood before and just in smaller circles / different name.
- bsenftner 5mo agoIn some ways, this is similar to use of a Control Net. I've been doing this same technique for a while, using only SVGs as the base image. Works well.
- jere 5mo agoVery impressive, simple, and reliable. I'm sure it will be picked up by image generation labs soon.
- sparuchuri 5mo agoThis hack definitely falls in the “duh, why didn’t I think of that” category of tricks, but glad to now have it next time imagegen comes up short
- manmal 5mo agoEven the original stable diffusion app had image 2 image. It just didn’t work as well. I‘m not sure why this is supposed to be novel.
- Finbel 5mo agoIt's not novel in the sense that nobody knew about img2img. It's novel in the sense that nobody thought of using img2img to solve this problem in this way.
- TeMPOraL 5mo agoIt's novel if you never played with img2img, including especially several forms of (text+img)2img. Or, if you never tried editing images by text prompt in recent multimodal LLMs. That said, I spent plenty of time doing both, and yet it would probably take me a while to arrive at this approach. For some reason, the "draw a sketch, have a model flesh it out" approach got bucketed with Stable Diffusion in my mind, and multimodal LLMs with "take detailed content, make targeted edits to it". So I'm glad the OP posted it.
- vunderba 5mo agoThey’re actually quite good at it. I’ve had a number of situations where I’ve wanted to re-render some of my older comics. You can basically tell any SOTA multimodal model (NB, GPT-Image-X) to treat them as storyboards and prompt for a specific style: newprint, crosshatching, monochromatic ink sketch, etc. Another thing I’ve gotten very used to doing is avoiding the “one-shot” approach. If I generate something and don’t like the results, I bring it into Krita, move things around, redraw some elements, and then send it back in with instructions to just clean it up (remove any smudges or imperfections). The state-of-the-art models can do an astonishing job with that workflow. https://imgpb.com/eGDJIb https://imgpb.com/eGDJIb
- manmal 5mo agoOk it might just be me then. I view Nvidia‘s DLSS as a similar thing. There was even this meme that video games will in the future only output basic geometry and the AI layer transforms it into stunning graphics.
- tracerbulletx 5mo agoIve been doing charts for slides like this for a while. Noticed html viz was super reliable, but I could style it with diffusion model. Its very useful for data viz.
- danpalmer 5mo agoI'm glad that we're making progress towards a deeper understanding of what LLMs are inherently good at and what they're inherently bad at (not to say incapable of doing, but stuff that is less likely to work due to fundamental limitations). There's similarity here with, for example, defining the architecture of software, but letting an LLM write the functions. Or asking an LLM to write you the SQL query for your data analysis, rather than asking it to do your data analysis for you. What I'd really like to see is a more well defined taxonomy of work and studies on which bits work well with LLMs and which don't. I understand some of this intuitively, but am still building my intuition, and I see people tripping up on this all the time.
- locknitpicker 5mo ago> There's similarity here with, for example, defining the architecture of software, but letting an LLM write the functions. Not so long ago, this was how early adopters of LLM coding assistants claimed was the right way to use them in coding tasks: prompt to draft the outline, and then prompt to implement each function. There were even a few posts in HN on blogposts showing off this approach with terms inspired in animation work.
- danpalmer 5mo agoI'm not necessarily suggesting always getting down to literally the function level, although I think that gives you excellent quality control, but having a code-level understanding is clearly an important factor.
- nullsanity 5mo ago[dead]
- Sammi 5mo agoIn short, LLMs are pretty great at working at a single level of abstraction at a time. You can go from the highest level and all the way down to the lowest level with LLMs, you just have to work at it iteratively one level at a time.
- 5mo ago
- gwern 5mo agotldr: do a standard img2img workflow where you lay out a skeleton or skeleton or low-res version, and then turn it into the final high-quality photorealistic version, instead of trying to zeroshot it purely from a text prompt.
- BobbyTables2 5mo agoHow is it that LLMs aren’t good at rendering the sequence of numbers but can reliably put the supplied pieces all in the right order?
- mk_stjames 5mo agoBecause the image generation is powered by a diffusion model that is only guided by the transformer model and still has somewhat vague spatial representation especially when it comes to coupling things like counting and complex positioning. But by using the LLM to generate code like an SVG graphic is made up of, and then using a rasterized image of that SVG as an input to the diffusion model, this takes place of the raw noise input and guides the denoising process of the diffusion model to put the numerical parts in the right spots. The LLM is putting the SVG in the right order because the code that drives the SVG is just that - code - and the numerical order is easily defined there, even if it has to follow something like a spiral. Edit: although LLMs now also may be using thinking modes with their feedback during generation to help with complex positioning when drawing something like an SVG, as I just asked claude to generate me one such spiral number SVG and it did so interactively via thinking, and the code generated is incredibly explicit with positions, so, that must help. But the underlaying idea to two-step SVG-to-diffusion model is the real key here.
- deleted 5mo ago[deleted]
- smusamashah 5mo agoThis is just img2img where first image with correct structure was generated by code.
- jasonjmcghee 5mo agoPretty much what the author said- just gave some context for the uninitiated
- philsnow 5mo agoRight, but you can use a different (codegen) model to make that code.
- vunderba 5mo agoYup, that’s exactly what this is. If you’ve been using generative models since the early Stable Diffusion days, it’s a pretty common (and useful!) technique: using a sketch (SVG, drawn, etc) as an ad-hoc "controlnet" to guide the generative model’s output. Example: In the past I'd use a similar approach to lay out architectural visualizations. If you wanted a couch, chair, or other furniture in a very specific location, you could use a tool like Poser to build a simple scene as an approximation of where you wanted the major "set pieces". From there, you could generate a depth map and feed that into the generative model, at the time SDXL, to guide where objects should be placed.
- nullc 5mo agoInpainting/guiding from a sketch is how I've always used diffusion models. I thought everyone did that, or at least everyone who wasn't just trying to get some arbitrary filler material without much care of what the output looked like.
- choeger 5mo agoTransformers are great translators. So, yeah, starting with structured output like SVG is probably the best way to start. It should be fairly trivial to fix any logic errors in the structured output, too.
- jeffrallen 5mo agoI wish the opposite was true: that when I tell Gemini I want "a diagram of X" that it immediately breaks out Python and mathplotlib, instead of wasting my time with Nano Banana.
- wg0 5mo agoHas anyone had good luck with making consistent game art and assets?
- SomaticPirate 5mo agoinb4 this technique is subsumed into the next MoE model release LLMs are evolving so fast I wouldn’t be surprised if this technique was not needed in <6 months
- rimliu 5mo agoLLMs are rather devolving at this point.
- krackers 5mo agoI don't think the MoE part has anything to do with it, but the current gen of multimoddal models can do thinking interleaved with autoregressive(?*) image-gen so it's probably not long before they bake this into the RL process, same way native thought obviated need for "think carefully step by step" prompts.
- Melamune 5mo agoI wondered why I was losing all passion for creating. These tips and tricks are part of the answer.
- xigoi 5mo agoThe standard objection: if the LLM is supposedly intelligent, why can’t it figure out on its own that this two-step process would achieve a better result?
- nine_k 5mo agoNobody asked it to!
- xigoi 5mo agoIf it’s asked to generate an image, it should to everything in its powers to make the image good.
- andruby 5mo ago> it should do everything in its powers That's a scary thought. Hey Claude, why haven't you finished yet? ... Because the human I'm holding hostage hasn't finished the drawing yet.
- lacksjoian 5mo agoLLMs have no concept of what makes the output "good". Or to put it another way, if the LLM generates an image with jumbled numbers it's because that was the most likely output, hence it was a "good" image according to its weights.
- cubefox 5mo agoPart of the problem is that it isn't the LLM making the image directly itself, it's the LLM repeatedly prompting edits for a separate edit diffusion model. The Gemini reasoning summary shows part of this. The style of some of the images makes it also clear that it uses an Imagen 4 derived diffusion model underneath.
- jstanley 5mo ago[flagged]
- 5mo ago
- nine_k 5mo agoIt's normal to first create a plan, then allow agents to write code. But it seems to be surprising for many to first create a draft / outline of a picture, then go for a final render.
- nottorp 5mo agoLLMs are like a box of chocolates...
- psychoslave 5mo agoA few months ago I tried to make Le-chat Mistral output a French poetry in Alexandrin (12 vowels). Disastrous at first. Then adding in specifications that each line had to also be transposed in IPA and each syllable counted, it went better. Still emotionally unrelatable, but definitely was providing something that match the specifications of there are explicit and systematically enforced through deterministitic means. For now I retain that LLM limitations are thus that they can't seize the ineffable and so untrustworthy they can only be employed under very clear and inescapable constraints or they will go awry just as sure as water is wet.
- dllu 5mo agoI was thinking about doing the opposite for the common task of "SVG of a pelican riding a bike". Obviously, directly spitting out the SVG is gonna be bad. But image gen can produce a really stunning photorealistic image easily. Probably a good way to get an LLM to produce a decent bike-pelican SVG is to generate an image first and then get the model to trace it into an SVG. After all, few human beings can generate SVG works of art by just typing out numbers into Notepad. At the core of it, we still rely on looking at it and thinking about it as an image.
- globular-toast 5mo agoWait, where did it get the "Sweet Path//Trail of treats" thing from in the SVG? It wasn't about sweets at that point. Something missing here, I think.
- elil17 5mo agoI wonder whether this could be used to fine-tune image models to provide better outputs. Something like this: 1. Algorithmically generate a underdrawing (e.g. place numbers and shapes randomly in the underdrawing) 2. Algorithmically generate a description of the underdrawing (e.g. for each shape, output text like "there is a square with the number three in the top left corner). You might fuzz this by having an LLM rewrite the descriptions in a variety of ways. 3. Generate a "ground truth" image using the underdrawing and an image+text-to-image model. 4. Use the generated description and the generated "ground truth" image as training data for a text-to-image model.
- hirako2000 5mo agoThat would complexity the architecture of a model, to solve a finite set of cases. That's an argument for specialised/fine tuned models though.
- slickytail 5mo ago[dead]
- vunderba 5mo agoThis is closer to a world model - kind of similar to how one might use a realistic or semi‑realistic simulation engine to model the environment like GTA in order to train a self-driving model.
- foxes 5mo agoI feel sorry for the recipient.
- cheekyant 5mo agoHas anyone built a platform which has image to image pipelines and lets you use prompt to SVG generation from SOTA LLMs?
- TeMPOraL 5mo agoComfyUI?
- Geonode 5mo agoWe've been doing this for a long time now, it's similar to using a depth map or a line drawing to control the silhouette.
- usagisushi 5mo agoYeah, I’ve used a similar technique to build a "pizza clock" before (where the number of slices corresponds to the hours).
- IdiotSavage 5mo ago> Transform this image into a photographed claymation diorama of assorted artisan chocolates and candies […] viewed from a low-angle Side note: whenever I read prompts for image generation, I notice very specific details which the model obviously ignored. Here the chocolates / candies in the last two images look anything but artisanal. They look very "sterile" and mass-produced. The viewing angle is also not accurate. Why do we even bother writing such elaborate prompts, when the model ignores most of it anyway?
- 8-prime 5mo agoI have noticed the same thing.The few times I wanted to use image generatation it always failed me in exactly these aspects. I always put if off as a lack of prompting skill on my end. Once you start to keep an eye out for these inconsistencies they turn out to be very common.
- ErroneousBosh 5mo agoI wonder how long it took to come up with all this? Because if I wanted a spiral of little "buttons" like the last one at the end (and they don't look very much like sweets) I'd be able to knock that out in Blender in an afternoon, and I'm not very good at Blender.
- HotHotLava 5mo agoI think you're vastly overestimating the average persons ability to use Blender if you can do that in an afternoon; just figuring out how to place a colored cube and the camera probably takes an afternoon if you pick up Blender for the first time.
- ErroneousBosh 5mo agoI guess I'm coming at it from having used Blender for an afternoon or so, and already knowing Python. If you were good at GLSL you could do it in that maybe. Someone somewhere is going to write something that directly draws it to a framebuffer in Brainfuck, you just know it, don't you?
- docheinestages 5mo agoAnd what happens if the model can't come up with a good enough SVG to begin with?
- utopiah 5mo agoLove the concluding note : it works, but not really. So LLM/GenAI crave. An entire article to show that it's nearly there, yet it's not, despite convoluted effort to make it just so on a very very niche example.
- Al-Khwarizmi 5mo agoBut if it works part of the time, it's useful. It's easy for a human to check that the numbers are correct, and if they aren't, just regenerate the image. Orders of magnitude easier than creating the image from scratch without the model.
- brentcrude 5mo ago[dead]
- petercooper 5mo agoThis seems analogous to how a human would do it accurately. If you asked an artist to paint stones in a large circular arrangement with the numbers in order in one shot, with no fixes or sketching allowed, it wouldn't be surprising to end up with problems in the arrangement.
- teiferer 5mo agoI hope this kind of stuff puts the idea to rest that we're close to actual AGI. Outsourcing this kind of basic stuff which a real intelligence would be able to do "internally" is a hack which works for this specific case but would prevent further generalizations of the task at hand. But I'm forseeing the opposite. This kind of tool use will soon be integrated and hidden such that people will eventully say "see we solved the problem that AI can't do 123+456, now we are really really close to AGI. Yeah no, with an AGI, it would have been the AGI itself that would have come up with needing at tool, building the tool and then using the tool. But that's not what LLMs are. They are statistical machines to predict tokens. They are very good at it, but that's not an AGI.
- globular-toast 5mo agoLike many AI things, it would have been considerably easier just to learn to edit images in GIMP or something. Instead of learning a valuable skill, you spent time working with a model that will be obsolete in a few months. Sunken cost fallacy, I guess.
- igtztorrero 5mo agoTask like this, are the cause IA is eating memory and cpu prices.
- oh_no 5mo agointeresting that GPT Image-2 managed to 2-shot this with thinking turned on, I didn't save a copy and it disappeared from my window but I first got a failure very similar to the one in the article, but it saw the issue and said it was going to use a reference image, after which it came out with https://i.imgur.com/hlWpQNT.jpeg https://i.imgur.com/hlWpQNT.jpeg
- npilk 5mo agoStill missing 49 - humans are safe, for now!
- kfarr 5mo agoI work on a platform 3dstreet.com that does “underdrawing” but in 3d space which image models also struggle with. Another company intangible.ai does this as well: low poly 3d then image to image model. It seems to be a very effective pattern. Curious if there are other examples out there. Or other names for this?
- barbazoo 5mo agoI've had a lot of success at work breaking down tasks that are supposed to be "done by the agent" into small LLM calls orchestrated deterministically via boring queues and messages. That's why this really resonates with me in a world where we're lured deep into the ecosystem by the model vendors. At the end of the day we can get so much done just by breaking down a problem into smaller problems.
- mncharity 5mo ago> deterministically But won't it be fun when we can cloud burst-parallel a grid/tiled sampling of multiple code implementations/architectures, and interactively explore navigate/blend-points-in the latent design space. Multiples embodying different trade-offs, styles, clarity vs performance, etc. Code as generative art. What might the software engineering equivalent of designer moodboards be?
- barbazoo 5mo agoI guess that would be cool.
- legalmoneytalk 5mo ago[flagged]
- sailorganymede 5mo agoI run an art training app and one of the features I am working on is AI generated critique (it's an app for training realism.) And I noticed if I gave the AI the original image, the tonal breakdown of the image the app generates as well as the user's drawing, it delivered critique far better and more targeted.