11 ms·
What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel
by blixt 2y ago
What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space.
Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on.
You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff like "change day to night", or "put a hat on him", and so forth.
I get the feeling these models are quite restricted in resolution, and that more work in this space will let us do really wild things such as ask a model to create an app step by step first completely in images, essentially designing the whole app with text and all, then writing the code to reproduce it. And it also means that a model can take over from a really good diffusion model, so even if the original generations are not good, it can continue "reasoning" on an external image.
Finally, once these models become faster, you can imagine a truly generative UI, where the model produces the next frame of the app you are using based on events sent to the LLM (which can do all the normal things like using tools, thinking, etc). However, I also believe that diffusion models can do some of this, in a much faster way.
- Mond_ 2y agoPretty sure the modern Gemini image models can already do token based image generation/editing and are significantly better and faster.
- famouswaffles 2y agoIt's faster but it's definitely not better than what's being showcased here. The quality of Flash 2 Image gens are generally pretty meh.
- blixt 2y agoYeah Gemini has had this for a few weeks, but much lower resolution. Not saying 4o is perfect, but my first few images with it are much more impressive than my first few images with Gemini.
- yieldcrv 2y agoweeks, ya'll, weeks!
- deleted 2y ago[deleted]
- jjbinx007 2y agoIt still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.
- tobr 2y agoAlso still seems to have a hard time consistently drawing pentagons. But at least it does some of the time, which is an improvement since last time I tried, when it would only ever draw hexagons.
- yusufozkan 2y agoAre you sure you are using the new 4o image generation? https://imgur.com/a/wGkBa0v https://imgur.com/a/wGkBa0v
- minimaxir 2y agoThat is an unexpectedly literal definition of "full glass".
- numpad0 2y agoExcept this is correct in this context. None of existing Diffusion models could, apparently.
- yusufozkan 2y agoGenerating an image of a completely full glass of wine has been one of the popular limitations of image generators, the reason being neural networks struggling to generalise outside of their training data (there are almost no pictures on the internet of a glass "full" of wine). It seems they implemented some reasoning over images to overcome that.
- kube-system 2y agoI wonder if that has changed recently since this has become a litmus test. Searching in my favorite search engine for "full glass of wine", without even scrolling, three of the images are of wine glasses filled to the brim.
- rafram 2y ago> Finally, once these models become faster, you can imagine a truly generative UI, where the model produces the next frame of the app you are using based on events sent to the LLM With current GPU technology, this system would need its own Dyson sphere.
- xg15 2y ago> What's important about this new type of image generation that's happening with tokens rather than with diffusion That sounds really interesting. Are there any write-ups how exactly this works?
- lyu07282 2y agoWould be interested to know as well. As far as I know there is no public information about how this works exactly. This is all I could find: > The system uses an autoregressive approach — generating images sequentially from left to right and top to bottom, similar to how text is written — rather than the diffusion model technique used by most image generators (like DALL-E) that create the entire image at once. Goh speculates that this technical difference could be what gives Images in ChatGPT better text rendering and binding capabilities. https://www.theverge.com/openai/635118/chatgpt-sora-ai-image-generation-chatgpt https://www.theverge.com/openai/635118/chatgpt-sora-ai-image...
- treis 2y agoI wonder how it'd work if the layers were more physical based. In other words something like rough 3d shape -> details -> color -> perspective -> lighting. Also wonder if you'd get better results in generating something like blender files and using its engine to render the result.
- astrange 2y agoDALL-E was an autoregressive encoder; it's 2 and 3 that used diffusion and were much less intelligent as a result.
- fpgaminer 2y agoThere are a few different approaches. Meta documents at least one approach quite well in one of their llama papers. The general gist is that you have some kind of adapter layers/model that can take an image and encode it into tokens. You then train the model on a dataset that has interleaved text and images. Could be webpages, where images occur in-between blocks of text, chat logs where people send text messages and images back and forth, etc. The LLM gets trained more-or-less like normal, predicting next token probabilities with minor adjustments for the image tokens depending on the exact architecture. Some approaches have the image generation be a separate "path" through the LLM, where a lot of weights are shared but some image token specific weights are activated. Some approaches do just next token prediction, others have the LLM predict the entire image at once. As for encoding-decoding, some research has used things as simple as Stable Diffusion's VAE to encode the image, split up the output, and do a simple projection into token space. Others have used raw pixels. But I think the more common approach is to have a dedicated model trained at the same time that learns to encode and decode images to and from token space. For the latter approach, this can be a simple model, or it can be a diffusion model. For encoding you do something like a ViT. For decoding you train a diffusion model conditioned on the tokens, throughout the training of the LLM. For the diffusion approach, you'd usually do post-training on the diffusion decoder to shrink down the number of diffusion steps needed. The real crutch of these models is the dataset. Pretraining on the internet is not bad, since there's often good correlation between the text and the images. But there's not really good instruction datasets for this. Like, "here's an image, draw it like a comic book" type stuff. Given OpenAI's approach in the past, they may have just bruteforced the dataset using lots of human workers. That seems to be the most likely approach anyway, since no public vision models are quite good enough to do extensive RL against. And as for OpenAI's architecture here, we can only speculate. The "loading from top to be from a blurry image" is either a direct result of their architecture or a gimmick to slow down requests. If the former, it means they are able to get a low resolution version of the image quickly, and then slowly generate the higher resolution "in order." Since it's top-to-bottom that implies token-by-token decoding. My _guess_ is that the LLM's image token predictions are only "good enough." So they have a small, quick decoder take those and generate a very low resolution base image. Then they run a stronger decoding model, likely a token-by-token diffusion model. It takes as condition the image tokens and the low resolution image, and diffuses the first patch of the image. Then it takes as condition the same plus the decoded patch, and diffuses the next patch. And so forth. A mixture of approaches like that allows the LLM to be truly multi-modal without the image tokens being too expensive, and the token-by-token diffusion approach helps offset memory cost of diffusing the whole image. I don't recall if I've seen token-by-token diffusion in a published paper, but it's feasible and is the best guess I have given the information we can see. EDIT: I should note, I've been "fooled" in the past by OpenAI's API. When o* models first came out, they all behaved as if the output were generated "all at once." There was no streaming, and in the chat client the response would just show up once reasoning was done. This led me to believe they were doing an approach where the reasoning model would generate a response and refine it as it reasoned. But that's clearly not the case, since they enabled streaming :P So take my guesses with a huge grain of salt.
- abossy 2y agoThat's very interesting. I would have assumed that 4o is internally using a single seed for the entire conversation, or something analogous to that, to control randomness across image generation requests. Can you share the technical name for this reasoning process so I could look up research about it?
- SpaceManNabs 2y agomultimodal chain of thought / generation of thought Nobody has really decided on a name. Also chain of thought is somewhat different from chain of thought reasoning so mb throw in multimodal chain of thought reasoning
- nine_k 2y agoIt also would mean that the model can correctly split the image into layers, or segments, matching the entities described. The low-res layers can then be fed to other image-processing models, which would enhance them and fill in missing small details. The result could be a good-quality animation, for instance, and the "character" layers can even potentially be reusable.
- SamBam 2y agoHmmm, I wanted to do that tic tac toe example, and it failed to create a 3x3 grid, instead creating a 5x5 (?) grid with two first moves marked. https://chatgpt.com/share/67e32d47-eac0-8011-9118-51b81756ec96 https://chatgpt.com/share/67e32d47-eac0-8011-9118-51b81756ec...
- nerder92 2y agoI tried to play it, and while the conversation is right the image is just all wrong
- M4v3R 2y agoTried it myself on the new model, worked out pretty well: https://chatgpt.com/share/67e34558-5244-8004-933a-23896c738bed https://chatgpt.com/share/67e34558-5244-8004-933a-23896c738b...
- jay_kyburz 2y agoI might just be a grumpy old man, but it really bugs me when the AI confidently says, "Here is your image, If you have any other requests, just let me know!". For a start the image is wrong, and also I know I can make more requests, because that what tools are for. Its like a passive aggressive suggestion that I made the AI go out of its way to do me a favor.
- doctoboggan 2y agoYour images say "Created with DALL-E", so you have not tried out the new model yet. I think they are gradually rolling it out.
- cchance 2y agoclick the ...'s and make sure you disable images (w/ dall-e) apparently the native just works if u enable the images... it switches to dall-e lol
- Taek 2y ago> What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. I do not think that this is correct. Prior to this release, 4o would generate images by calling out to a fully external model (DALL-E). After this release, 4o generates images by calling out to a multi-modal model that was trained alongside it. You can ask 4o about this yourself. Here's what it said to me: "So while I’m deeply multimodal in cognition (understanding and coordinating text + image), image generation is handled by a linked latent diffusion model, not an end-to-end token-unified architecture."
- CooCooCaCha 2y agoModels are famously good at understanding themselves.
- uh_uh 2y agoI hope you're joking. Sometimes they don't even know which company developed them. E.g. DeepSeek was claiming it was developed by OpenAI.
- SketchySeaBeast 2y agoWell, that one seems to be true, from a certain point of view.
- notduncansmith 2y agoI hope you’re joking :)
- enceladus06 2y agoI have asked GPT if it is using the 4o or 4.5 model multiple times in voice mode e.g. "Which model are you using?". It has said that it is using 4.5 when it is actually using 4o.
- rickyhatespeas 2y agoYou're incorrect. 4o was not trained on knowledge of itself so literally can't tell you that. What 4o is doing isn't even new either, Gemini 2.0 has the same capability.
- sureIy 2y ago> truly generative UI, where the model produces the next frame of the app Please sir step away from the keyboard now! That is an absurd proposition and I hope I never get to use an app that dreams of the next frame. Apps are buggy as they are, I don't need every single action to be interpreted by LLM. An existing example of this is that AI Minecraft demo and it's a literal nightmare.
- blixt 2y agoThis argument could be made for every level of abstraction we've added to software so far... yet here we are commenting about it from our buggy apps!
- outworlder 2y agoYeah, but the abstractions have been useful so far. The main advantage of our current buggy apps is that if it is buggy today, it will be exactly as buggy tomorrow. Conversely, if it is not currently buggy, it will behave the same way tomorrow. I don't want an app that either works or does not work depending on the RNG seed, prompt and even data that's fed to it. That's even ignoring all the absurd computing power that would be required.
- blixt 2y agoStill sounds a bit like we've seen it all already – dynamic linking introduced a lot of ways for software that wasn't buggy today to become buggy tomorrow. And Chrome uses an absurd amount of computing power (its bare minimum is many multiples of what was once a top-of-the-line, expensive PC). I think these arguments would've been valid a decade ago for a lot of things we use today. And I'm not saying the classical software way of things needs to go away or even diminish, but I do think there are unique human-computer interactions to be had when the "VM" is in fact a deep neural network with very strong intelligence capabilities, and the input/output is essentially keyboard & mouse / video+audio.
- jychang 2y agoYou're just describing calling a customer service phone line in India.
- jacobsenscott 2y ago> writing the code to reproduce it I'm super excited for all the free money and data our new AI written apps will be giving away.
- snickell 2y ago> truly generative UI, where the model produces the next frame of the app I built this exact thing last month, demo: https://universal.oroborus.org https://universal.oroborus.org (not viable on phone for this demo, fine on tablet or computer) Also see discussion and code at: http://github.com/snickell/universal http://github.com/snickell/universal I wasn't really planning to share/release it today, but, heck, why not. I started with bitmap-style generative image models, but because they are still pretty bad at text (even this, although it’s dramatically better), for early-2025 it’s generating vector graphics instead. Each frame is an LLM response, either as an svg or static html/css. But all computation and transformation is done by the LLM. No code/js as an intermediary. You click, it tells the LLM where you clicked, the LLM hallucinates the next frame as another svg/static-html. If it ran 50x faster it’d be an absolutely jaw dropping demo. Unlike "LLMs write code", this has depth. Like all programming, the "LLMs write code" model requires the programmer or LLM to anticipate every condition in advance. This makes LLM written "vibe coded" apps either gigantic (and the llm falls apart) or shallow. In contrast, as you use universal, you can add or invent features ranging from small to big, and it will fill in the blanks on demand, fairly intelligently. If you don't like what it did, you can critique it, and the next frame improves. Its agonizingly slow in 2025, but much smarter and in weird ways less error prone than using the LLM to generate code that you then run: just run computation via the LLM itself. You can build pretty unbelievable things (with hallucinated state, granted) with a few descriptive sentences, far exceeding the capabilities you can “vibe code” with the description. And it never gets lost in its rats nest of self generated garbage code because… there is no code to in. Code is medium with a surprisingly strong grain. This demo is slow, but SO much more flexible and personally adaptable than anything I’ve used where the logic is implemented cia a programming language. I don’t love this as a programmer, but my own use of the demo makes me confident that programming languages as a category will have a shelf life if LLM hardware gets fast, cheap and energy efficient. I suspect LLMs will generate not programming language code, but direct wasm or just machine code on the fly for things that need faster traction than they can draw a frame, but core logic will move out of programming languages (not even llm written code). Maybe similar to the way we bind to low level fast languages but a huge percentage of “business” logic is written in relatively slower languages. FYI, I may not be able to afford the credits if too many people visit, I put a a $1000 of credits on this, we'll see if that lasts. This is claude 3.7, I tried everything else, a claude had the visual intelligence today. IMO this is a much more compelling glance at the future than coding models. Unfortunately, generating an SVG per click is pricey, each click/frame costs me about $0.05. I’ll fund this as far as I can so folks can play with it. Anthropic? You there? Wanna throw some credits at an open source project doing something that literally only works on claude today? Not just better, but “only Claude 3.7 can show this future today?”. I’d love for lots more people to see the demo, but I really could use an in-kind credit donation to make this viable. If anyone at anthropic is inspired and wants to hook me up: snickell@alumni.stanford.edu. Very happy to rep Claude 3.7 even more than I already do. I think it’s great advertising for Claude. I believe the reason Claude seems to do SO much better at this task is, one it shows far greater spatial intelligence, and two, I distract they are the only state of the art model intentionally training on SVG.
- canjobear 2y agoWrt reasoning I’ll believe it when I see it. I just tried several variants of “Generate an image of a chess board in which white has played three great moves and black has played two bad moves.” Results are totally nonsensical as always.
- DeathArrow 2y ago>You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff like "change day to night", or "put a hat on him", and so forth. You can do that with diffusion, too. Just lock the parameters in ComfyUi.
- blixt 2y agoYeah I wasn’t very imaginative in my examples, with 4o you can also perform transformations like “rotate the camera 10 degrees to the left” which would be hard without a specialized model. Basically you can run arbitrary functions on the exact image contents but in latent space.
- DeathArrow 2y agoIt's doable with diffusion, too.
- echelon 2y agoI'm incredibly deep in the image / video / diffusion / comfy space. I've read the papers, written controlnets, modified architectures, pretrained, finetuned, etc. All that to say that I've been playing with 4o for the past day, and my opinions on the space have changed dramatically. 4o is a game changer. It's clearly imperfect, but its operating modalities are clearly superior to everything else we have seen. Have you seen (or better yet, played with) the whiteboard examples? Or the examples of it taking characters out of reflections and manipulating them? The prompt adherence, text layout, and composing capabilities are unreal to the point this looks like it completely obsoletes inpainting and outpainting. I'm beginning to think this even obsoletes ComfyUI and the whole space of open source tools once the model improves. Natural language might be able to accomplish everything outside of fine adjustments, but if you can also supply the model with reference images and have it understand them, then it can do basically everything. I haven't bumped into anything that makes me question this yet. They just need to bump the speed and the quality a little. They're back at the top of image gen again. I'm hoping the Chinese or another US company releases an open model capable of these behaviors. Because otherwise OpenAI is going to take this ball and run far ahead with it.
- hansmayer 2y agoA lot of convoluted explanations about something we don't even know if it really works all the time. I feel like in the third year of LLM-Hype and after reminde-me-how-many billions of dollars burned, we should by now not have to imagine what 'might happen' down to road, it should have been happening already. The use-case you are describing, sure sounds very interesting, until I remember asking asked Copilot for a simple scaffolding structure in React, and it spat out something which lacked half of imports and proper visual alignments. A few years ago I was excited about the possibility of removing all the scaffolding and templating work so we can write the cool parts, but they cannot even do that right. It's actually a step back compared to automatic code generators of the past, because those at least produced reproducible results every single time you used them. But hey sure, the next generation of "AI" (it's not really AI) will probably solve it.
- mckngbrd 2y agoConsider using a better AI IDE platform than Copilot ... cursor, windsurf, cline, all great options that do much better than what you're describing. The underlying LLM capabilities also have advanced quite a bit in the past year.
- hansmayer 2y agoWell I do not really use it that much to actually care, and don't really depend on AI, thankfully. If they did not mess up the google search, we wouldnt even need that crap at all. But that's not the main point. Even if I switched to cursor or windsurf - aren't they all using one of the same LLMs? (ChatGPT,Claude, whatever..). The issue is that the underlying general approach will never be accurate enough. There is a reason most of successful technologies lift off quickly and those not successful also die very quickly. This is a tech propped up by a lot of VC money for now, but at some point, even the richest of the rich VCs will have trouble explaining spending 500B dollars in total, to get something like 15B revenue (not even profit). And don't even get me started on Altman's trillion-fantasies...
- mckngbrd 2y ago
- belter 2y agoIs it able to break the usual failure modes of these models, that all clocks are at 10 min past two, or they can't produce images of people drawing with the left hand?
- lyu07282 2y agoIn my tests no, that's still not possible with the model unfortunately, but it feels like you have way more control with prompting over any previous model (stable diffusion/midjourney).