16 ms·
This is the gpt 4 moment for image editing models. Nano banana aka gemini 2.5 flash is insanely good. It made a 171 elo point jump in lmarena! Just search nano
by fariszr 1y ago
This is the gpt 4 moment for image editing models.
Nano banana aka gemini 2.5 flash is insanely good.
It made a 171 elo point jump in lmarena!
Just search nano banana on Twitter to see the crazy results.
An example.
https://x.com/D_studioproject/status/1958019251178267111 https://x.com/D_studioproject/status/1958019251178267111
- ceroxylon 1y agoIt seems like every combination of "nano banana" is registered as a domain with their own unique UI for image generation... are these all middle actors playing credit arbitrage using a popular model name?
- bonoboTP 1y agoI'd assume they are just fake, take your money and use a different model under the hood. Because they already existed before the public release. I doubt that their backend rolled the dice on LMArena until nano-banana popped up. And that was the only way to use it until today.
- ceroxylon 1y agoAgreed, I didn't mean to imply that they were even attempting to run the actual nano banana, even through LMarena. There is a whole spectrum of potential sketchiness to explore with these, since I see a few "sign in with Google" buttons that remind me of phishing landing pages.
- vunderba 1y agoThey're almost all scams. Nano banana AI image generator sites were showing up when this model was still only available in LM Arena.
- koakuma-chan 1y agoWhy is it called nano banana?
- ehsankia 1y agoBefore a model is announced, they use codenames on the arenas. If you look online, you can see people posting about new secret models and people trying to guess whose model it is.
- mvdtnz 1y agoWhat are "the arenas"?
- patates 1y agoBlind rating battlegrounds, one is https://lmarena.ai/ https://lmarena.ai/ (first google result)
- kstenerud 1y agoI don't quite get what this is? I asked the AI on the site "What is imarena.ai?" and it just gave some hallucinated answer that made no sense.
- adventured 1y agoPeople vote on the performance of AI, generating ranking boards.
- kstenerud 1y agoAh, that was the missing piece of information! Thanks!
- Jensson 1y agoEngineers often have silly project names internally, then some marketing team rewrites the name for public release.
- ZephyrBlu 1y agoI'm pretty sure it's because an image of a banana under a microscope generated by the model went super viral
- dcre 1y agoAlarming hands on the third one: it can't decide which way they're facing. But Gemini didn't introduce that, it's there in the base image.
- 725686 1y agoYes, the base image's hands are creepy.
- meatmanek 1y agoI noticed the AI pattern on the sunglasses first. I guess all of the source images are AI-generated? In a sense, that makes the result slightly less impressive -- is it going to be as faithful to the original image when the input isn't already a highly likely output for an AI model? Were the input images generated with the same model that's being used to manipulate them?
- dcre 1y agoIt doesn't seem to matter: people have posted tons of examples on social media of non-AI base images that it was equally able to hold steady while making edits.
- echelon 1y ago> This is the gpt 4 moment for image editing models. No it's not. We've had rich editing capabilities since gpt-image-1, this is just faster and looks better than the (endearingly? called) "piss filter". Flux Kontext, SeedEdit, and Qwen Edit are all also image editing models that are robustly capable. Qwen Edit especially. Flux Kontext and Qwen are also possible to fine tune and run locally. Qwen (and its video gen sister Wan) are also Apache licensed. It's hard not to cheer Alibaba on given how open they are compared to their competitors. We've left the days of Dall-E, Stable Diffusion, and Midjourney of "prompt-only" text to image generation. It's also looking like tools like ComfyUI are less and less necessary as those capabilities are moving into the model layer itself.
- bsenftner 1y agoI'm totally with you. Dismayed by all these fanbois.
- raincole 1y agoIn other words, this is the gpt 4 moment for image editing models. Gpt4 isn't "fundamentally different" from gpt3.5. It's just better. That's the exact point the parent commenter was trying to make.
- retinaros 1y agodid you see the generated pic demis posted on X? it looks like slop from 2 years ago. https://x.com/demishassabis/status/1960355658059891018 https://x.com/demishassabis/status/1960355658059891018
- raincole 1y agoI've tested it on Google AI Studio since it's available to me (which is just a few hours so take it with a grain of salt). The prompt comprehension is uncannily good. My test is going to https://unsplash.com/s/photos/random https://unsplash.com/s/photos/random and pick two random images, send them both and "integrate the subject from the second image into the first image" as the prompt. I think Gemini 2.5 is doing far better than ChatGPT (admittedly ChatGPT was the trailblazer on this path). FluxKontext seems unable to do that at all. Not sure if I were using it wrong, but it always only considers one image at a time for me. Edit: Honestly it might not be the 'gpt4 moment." It's better at combining multiple images, but now I don't think it's better at understanding elaborated text prompt than ChatGPT.
- rplnt 1y agoOh no, even more mis-scaled product images.
- qingcharles 1y agoI've been testing it for several weeks. It can produce results that are truly epic, but it's still a case of rerolling the prompt a dozen times to get an image you can use. It's not God. It's definitely an enormous step though, and totally SOTA.
- spaceman_2020 1y agoIf you compare to the amount of effort required in Photoshop to achieve the same results, still a vast improvement
- qingcharles 1y agoI work in Photoshop all day, and I 100% agree. Also, I just retried a task that wouldn't work last night on nano-banana and it worked first time on the released model, so I'm wondering if there were some changes to the released version?
- spaceman_2020 1y agoWe had an exhibition some time back where I used AI to generate the posters for our product. This is a side project and not something we do seriously, but the results were outstanding - better than what the majority of much bigger exhibitors had. It took me a LOT of time to get things right, but if I was to get an actual studio to make those images, it would have cost me a thousands of dollars
- Bombthecat 1y agoYeah, played around with it, it created an amazing poster for starfinder ttrpg ( something like DND) with specifies who looked really! Good. Usually stuff likes this fails hard, since there isn't much training data of unique fantasy creatures. But flash 2.5? Worked! It did it, crazy stuff
- Bombthecat 1y agoHow many times did you tried? I uploaded a black and white photo and let it colourize, something like 20 percent were still black and white.
- 93po 1y agoCompletely agree - I make logos for my github projects for fun, and the last time I tried SOTA image generation for logos, it was consistently ignoring instructions and not doing anything close to what i was asking for. Google's new release today did it near flawlessly, exactly how I wanted it, in a single prompt. A couple more prompts for tweaking (centering it, rotating it slightly) got it perfect. This is awesome.
- summerlight 1y agoI wonder how the creative workflow looks like when this kind of models are natively integrated into digital image tools. Imagine fine-grained controls on each layer and their composition with the semantic understanding on the full picture.
- torginus 1y agoAnother nitpick - the pink puffer jacket that got edited into the picture is not the same as the one in the reference image - it's very similar but if I were to use this model for product placement, or cared about these sort of details, I'd definitely have issues with this.
- drmath 1y agoEven in the just-photoshop-not-ai days product photos had become pretty unreliable as a means of understanding what you're buying. Of course it's much worse now.
- ethbr1 1y agoNote: Please understand that monitor may color different. If image does not match product received then kindly your monitor calibration. Seller not responsible. /ebay&amazon
- wiz21c 1y agolook at the bottom of the sleeves, they don't match. the bottom of the jacket doesn't match either. I didn't see it at first sight but it certainly is not the same jacket. If you use that as an advertisement, people can sue you for lying about the product.
- ivape 1y agoRegardless, it seems Google is on the frontier of every type of model and robotics (cars). It’s nutty how we forget what a intellectual juggernaut they are.
- fariszr 1y agoTool use and sycophancy are still big issues in gemini 2.5 models.
- fHr 1y agocope
- torginus 1y agoNo, it's not really that much of an improvement. Once you start coming up with specific tasks, it fails just like the others.
- hapticmonkey 1y agoBefore AI, people complained that Google was taking world class engineering talent and using it for little more than selling people ads. But look at that example. With this new frontier of AI, that world class engineering talent can finally be put to use…for product placement. We’ve come so far.
- torginus 1y agoI am pretty sure a lot of said engineering talent isn't actually contributing to AI but doing other stuff
- vineyardmike 1y ago> finally be put to use…for product placement. Did you think that Google would just casually allow their business to be disrupted without using the technology to improve the business and also protecting their revenue? Both Meta and Google have indicated that they see Generative AI as a way to vertically integrate within the ad space, disrupting marketing teams, copyrighters, and other jobs who monitor or improve ad performance. Also FWIW, I would suspect that the majority of Google engineers don't work on an ad system, and probably don't even work on a profitable product line.
- johnfn 1y agoOh come on - you have this incredible technology at your disposal and all you can think to use it for is product placement?
- polishdude20 1y agoThe fingernails on one of them. Ohhh nooo
- ethbr1 1y agoImage genai made me realize just how inattentive to detail a lot of people are.
- goosejuice 1y agoYet it's failed spectacularly at almost everything I've given it.
- Viaya 1y ago[dead]
- Viaya 1y ago[dead]
- mooncakes_ooohh 1y agoBe gone scammer
- r33b33 1y agonano banana is good, but not insanely good
- littlestymaar 1y ago> An example. https://x.com/D_studioproject/status/1958019251178267111 https://x.com/D_studioproject/status/1958019251178267111 “Nano banana” is probably good, given its score on the leaderboard, but the examples you show don't seem particularly impressive, it looks like what Flux Kontext or Qwen Image do well already.