9 ms·
Are you sure you are using the new 4o image generation? https://imgur.com/a/wGkBa0v https://imgur.com/a/wGkBa0v
by yusufozkan 2y ago
Are you sure you are using the new 4o image generation?
https://imgur.com/a/wGkBa0v https://imgur.com/a/wGkBa0v
- minimaxir 2y agoThat is an unexpectedly literal definition of "full glass".
- numpad0 2y agoExcept this is correct in this context. None of existing Diffusion models could, apparently.
- yusufozkan 2y agoGenerating an image of a completely full glass of wine has been one of the popular limitations of image generators, the reason being neural networks struggling to generalise outside of their training data (there are almost no pictures on the internet of a glass "full" of wine). It seems they implemented some reasoning over images to overcome that.
- kube-system 2y agoI wonder if that has changed recently since this has become a litmus test. Searching in my favorite search engine for "full glass of wine", without even scrolling, three of the images are of wine glasses filled to the brim.
- Loeffelmann 2y agoThat's the point. With the old models they all failed to produce a wine glass that is completley to the brim full. Because you can't find that a lot in the data they used for training.
- colecut 2y agoImagine if they just actually trained the model on a bunch of photographs of a full glass of wine, knowing of this litmus test
- HelloImSteven 2y agoEven if they did, I’d assume the association of “full” and this correct representation would benefit other areas of the model. I.e., there could (/should?) be general improvement for prompts where objects have unusual adjectives. So maybe training for litmus tests isn’t the worst strategy in the absence of another entire internet of training data…
- nefarious_ends 2y agoimagine!
- gorkish 2y agoI obviously have no idea if they added real or synthetic data to the training set specifically regarding the full-to-the-brim wineglass test, but I fully expect that this prompt is now compromised in the sense that because it is being discussed in the public sphere, it's has inherently become part of the test suite. Remember the old internet adage that the fastest way to get a correct answer online is to post an incorrect one? I'm not entirely convinced this type of iterative gap finding and filling is really much different than natural human learning behavior.
- vlovich123 2y agoHumans don’t train on the entire contents of the Internet, so i’d wager that they do learn differently
- sayamqazi 2y agoI think there is a critical aspect of human visual learning which machine leanring cant replicate because it is prohibitively expensive. When we look at things as children we are not just looking at a single snapshot. When you stare at an object for a few seconds you have practically injested hundreds of slightly variated images of that object. This gets even more interesting when you take into account real world is moving all the time, so you are seeing so many things from so many angles. This is simply undoable with compute.
- jorvi 2y agoThe old models were doing it correct also. There is no one correct way to interpert 'full'. If you go to a wine bar and ask for a full glass of wine, they'll probably interpert that as a double. But you could also interpert it the way a friend would at home, which is about 2-3cm from the rim. Personally I would call a glass of wine filled to the brim 'overfilled', not 'full'.
- drdeca 2y agoPeople were telling the models explicitly to fill it to the brim, and the models were still producing images where it was filled to approximately the half-way point.
- kalleboo 2y agoI think you're missing the context everyone else has - this video is where the "AI can't draw a full glass of wine" meme got traction https://www.youtube.com/watch?v=160F8F8mXlo https://www.youtube.com/watch?v=160F8F8mXlo The prompts (some generated by ChatGPT itself, since it's instructing DALL-E behind the scenes) include phrases like "full to the brim" and "almost spilling over" that are not up to interpretation at all.
- sejje 2y agoI did coax the old models into doing it once (dall-e) but it was like a fun exercise in prompting. They definitely didn't want to.
- yusufozkan 2y agoThis is another cool example from their blog https://imgur.com/a/Svfuuf5 https://imgur.com/a/Svfuuf5
- Imustaskforhelp 2y agoLooks amazing,can you please also create a unconventional image like the clock at 2:35 , I tried it something like this with gemini when some redditor asked it and it failed so wondering if 4o does do it
- Workaccount2 2y agoI tried and while the clock it generated was very well done and high quality, it showed the time as the analog clock default of 10:10.
- lyu07282 2y agoThe problem now is we don't know if people mistake dall-e for the new multimodal gpt4o output, they really should've made that clearer.
- cmorgan31 2y agoI’m using 4o and it gets time wrong a decent chunk but doesn’t get anything else in the prompt incorrect. I asked for the clock to be 4:30 but got 10:10. OpenAI pro account.
- Imustaskforhelp 2y agoShouldn't reasoning make the clock work though. Why does it sound like this isn't reasoning on images directly but rather just dall e as some other comment said , I will type the name of the person here (coder543)
- CSMastermind 2y agoI tried and it failed repeatedly (like actual error messages): > It looks like there was an error when trying to generate the updated image of the clock showing 5:03. I wasn’t able to create it. If you’d like, you can try again by rephrasing or repeating the request. A few times it did generate an image but it never showed the right time. It would frequently show 10:10 for instance.
- deleted 2y ago[deleted]
- stevesearer 2y agoCan you do this with the prompt of a cow jumping over the moon? I can’t ever seem to get it to make the cow appear to be above the moon. Always literally covering it or to the side etc.
- michaelt 2y agohttps://chatgpt.com/share/67e31a31-3d44-8011-994e-b7f8af7694d5 https://chatgpt.com/share/67e31a31-3d44-8011-994e-b7f8af7694... got it on the second try.
- coder543 2y agoTo be clear, that is DALL-E, not 4o image generation. (You can see the prompt that 4o generated to give to DALL-E.)
- dimitri-vs 2y agoHere you go: https://imgur.com/a/QJlj4I9 https://imgur.com/a/QJlj4I9