13 ms·
Nano Banana 2: Google's latest AI image generation model
- deleted 7mo ago[deleted]
- aliljet 7mo agoI really really want to see how these images are starting to form into videos. The stills are clearly getting better and better, but what about when you need the stills to organically conform to a keyed script?
- vessenes 7mo agothe workflow right now would be to take this images, make a sequence of them for key "shots" and send them to an I2V model. LTX-2 is the model the r/stablediffusion folks are playing with right now, but there are a fair few.
- progbits 7mo agoI'm seeing more and more AI video memes and they are getting really good. Still just bunch of short clips, long shots are not working well enough, but typical Hollywood movies have few second cuts anyway so this is almost good enough to make a marvel fanfic.
- Mizza 7mo agoCheck out Seedance 2: https://seed.bytedance.com/en/seedance2_0 https://seed.bytedance.com/en/seedance2_0 Nano Banana was technically impressive the first time, but after Seedance it's not really. It's all just an internet pollution machine anyway.
- sync 7mo agoDid gemini-2.5-flash-image get an upgrade as well? I just got the following, which is fascinating, and not something I've seen before: > I'm sorry, but I cannot fulfill your request as it contains conflicting instructions. You asked me to include the self-carved markings on the character's right wrist and to show him clutching his electromancy focus, but you also explicitly stated, "Do NOT include any props, weapons, or objects in the character's hands - hands should be empty." This contradiction prevents me from generating the image as requested. My prompts are automated (e.g. I'm not writing them) and definitely have contained conflicting instructions in the past. A quick google search on that error doesn't reveal anything either
- sorenjan 7mo agoIs this a distillation of Nano Banana Pro?
- meetpateltech 7mo agoGemini 3.1 Flash Image is based on Gemini 3 Flash. source: https://deepmind.google/models/model-cards/gemini-3-1-flash-image/ https://deepmind.google/models/model-cards/gemini-3-1-flash-...
- dgtlanml2 7mo agoWow the article narration with Umbriel is silent after the 6 second mark.
- vessenes 7mo agoInteresting they get to rev this with the release of a new flash model. I'm speculating part of the distil pipeline includes the image gen stuff; that seems like internal tooling that will pay dividends over time, if true. New frontier model -> automatic new image model. Even if it's just incremental updates, it's good for both the product cadence and compounding improvements.
- WarmWash 7mo agoThe confusion here is dense, 3.1 Flash Image is not 3.1 Flash. The banana models (image) are a different than the mainline models, but the confusingly leverage the same naming scheme.
- NitpickLawyer 7mo ago> the distil pipeline I don't have inside info, but everything we've seen about gemini3.0 makes me think they aren't doing distillation for their models. They are likely training different arch/sizes in parallel. Gemini 3.0-flash was better than 3.0-pro on a bunch of tasks. That shouldn't happen with distillation. So my guess is that they are working in parallel, on different arches, and try out stuff on -flash first (since they're smaller and faster to train) and then apply the learnings to -pro training runs. (same thing kinda happened with 2.5-flash that got better upgrades than 2.5-pro at various points last year). Ofc I might be wrong, but that's my guess right now.
- vessenes 7mo agoInteresting. Whatever they are doing it's a bit different than Anthropic and oAI, which is good for the consumer. I'm curious about their ML Ops internally; would be fascinating to learn more.
- minimaxir 7mo agoGoogle updated it early in AI Studio so I've been experimenting: - Base pricing for a 1024x1024 image is almost 1.6x what normal Nano Banana is ($0.067 vs. $0.039), however you can now get a 512x512 image for cheaper, or a 4k image for cheaper than four 1k images: https://ai.google.dev/gemini-api/docs/pricing#gemini-3.1-flash-image-preview https://ai.google.dev/gemini-api/docs/pricing#gemini-3.1-fla... - Thinking is now configurable between `Minimal` and `High` (was not the case with Nano Banana Pro) - Safety of the model appears to be increased so typical copyright infringing/NSFW content is difficult to generate (it refused to let me generate cartoon characters having taken psychedelics) - Generation speed is really slow (2-3min per image) but that may be due to load. - Prompt adherence to my trickier prompts for Nano Banana Pro (https://minimaxir.com/2025/12/nano-banana-pro/ https://minimaxir.com/2025/12/nano-banana-pro/) is much worse, unsurprisingly. For example I asked it to make a 5x2 grid with 10 given inputs and it keeps making 4x3 grids with duplicate inputs. However, I am skeptical with their marquee feature: image search. Anyone who has used Nano Banana Pro for awhile knows that it will strongly overfit on any input images by copy/pasting the subject without changes which is bad for creativity, and I suspect this implementation appears the same. Additionally I have a test prompt which exploits the January 2025 knowledge cutoff: Generate a photo of the KPop Demon Hunters performing a concert at Golden Gate Park in their concert outfits. That still fails even with Grounding with Google Search and Image Search enabled, and more charitable variants of the prompt. tl;dr the example images (https://deepmind.google/models/gemini-image/flash/ https://deepmind.google/models/gemini-image/flash/) seem similar to Nano Banana Pro which is indeed a big quality improvement but even relative to base Nano Banana it's unclear if it justifies a "2" subtitle especially given the increased cost.
- pietz 7mo agoI'm officially done with the Nano Banana name. It was fun, but can we go back just calling it Gemini Image?
- bonoboTP 7mo agoName recognition has big value. People remember what an advancement the first banana was. Nowadays it's no longer so unique, ChatGPT's and Grok's image editors are also strong.
- PunchTornado 7mo agoI really like it. Nano banana is like the best product name in AI.
- meowface 7mo agoHow does it compare to Nano Banana Pro?
- riteshyadav02 7mo ago[dead]
- yakattak 7mo agoI think this tech is cool, from an engineering perspective. I’m trying to figure out if there’s any justification for using it in a business world outside of: “We don’t want to pay an artist.” You can argue things like code generation are an extension of the engineer wielding it. Image generation just seems like a net negative overall if it’s used at scale. Edit: By scale, I mean large corporations putting content in front of millions. I understand the appeal for smaller businesses where they probably weren’t going to pay an artist anyway.
- bonoboTP 7mo agoDrafting, iteration, mockups. Quite useful during ideation.
- yakattak 7mo agoAll things traditionally done by artists or artist adjacent roles. I can understand at an individual level, say for a solo gamedev who wasn’t going to pay an artist anyway. That’s not at scale though. Larian Studios most recently was under fire for this [1]. Like I can see a director going “what would X look like?” and then speeding over to the concept artists for a proper rendition if they liked it. I don’t think this is at scale though. Any large business is just going to get rid of the concept artists. [1]: https://www.pcgamer.com/games/rpg/baldurs-gate-3-developer-larian-defends-itself-as-fans-react-to-generative-ai-use-im-not-entirely-sure-we-are-the-ideal-target-for-the-level-of-scorn/ https://www.pcgamer.com/games/rpg/baldurs-gate-3-developer-l...
- bonoboTP 7mo agoThere are many places in general office work where you need some kind of graphics. Slides, reports, info graphics, dataviz. Or academic papers. Some are just illustrations, like a fancy clipart or stock photos, some are drafts for a proper tikz or svg or something that you then redo in draw.io etc. There is much more use for graphics than the use cases where people would ever even consider hiring an actual artist. I've seen good results for iterating on eg model architecture figures quickly between PhD students and supervisors, faster than dragging boxes around and fiddling with tikz. Obviously you don't simply paste the result into the paper. You redo it but it's a good discussion basis. That's for info graphics stuff. But the same can apply to creative stuff, like an event poster, an invitation card to your wedding, storyboards, mood boards, DIY interior design, outfit planning etc etc
- nickandbro 7mo agoThese image gen models are getting so advanced and life like that increasingly the general public are being duped into believing AI images are actually real (ex Facebook food images or fake OF models). Don't get me wrong I will enjoy the benefits of using this model for expressing myself better than ever before, but can't help feeling there's something also very insidious about these models too.
- WarmWash 7mo agoIt's more likely than not that every single person who uses the internet has viewed an AI image and taken it as real by now. The obvious ones stand out, but there are so many that are indiscernible without spending lots of time digging through it. Even then there are ones that you can at best guess it's maybe AI gen.
- versk 7mo agoAt the point now where basically any photo that isn't shared by someone I trust or a reputable news organisation is essentially unverifiable as being real or not The positive aspect of this advance is that I've basically stopped using social media because of the creeping sense that everything is slop
- yieldcrv 7mo agopeople only notice when they are prompted to look for AI or scrutinize AI a lot of these accounts mix old clips with new AI clips or tag onto something emotional like a fake Epstein file image with your favorite politician, and pointing out its AI has people thinking you’re deflecting because you support the politician Meanwhile the engagement farmer is completely exempt from scrutiny Its fascinating how fast and unexpected the direction goes
- tokai 7mo agoMaybe not an actual argument for anything, but even before these image models everyone that used the internet had seen a doctored image they believed to be real. There was a reason that 'i can tell by the pixels' was a meme.
- 7mo ago
- nightski 7mo ago[flagged]
- DalasNoin 7mo agoWhy does SynthID make it worthless? it helps other platforms detect this as ai?
- zardo 7mo agoIf the value is in deception.
- csjh 7mo agoWhat’s the downside of SynthID?
- estearum 7mo agoQuite telling that you think a technology that merely prevents you from passing off an AI-generated image as not-AI-generated makes the model "worthless." Good! That's the point! Whatever amazing use case you had in mind is bad and I'm glad SynthID (apparently) makes it impossible.
- nightski 7mo agoActually no it just makes me use a different model. My uses are not nefarious at all, although it's fine for you to assume so. There are real, legitimate reasons why SynthId is actively harmful that do not involve deceiving or manipulating people at all. SynthId is just a stain on legitimate AI users. People who want to deceive or manipulate are not using Google models anyways. They are going to use a model without safety rails (which is not what I am advocating for per se, just that SynthId is an awful solution). It actually reeks of Google, since it's a technical solution to a people problem. Google doesn't seem to understand people.
- estearum 7mo ago> There are real, legitimate reasons why SynthId is actively harmful that do not involve deceiving or manipulating people at all I am legitimately curious: can you name some? > Actually no it just makes me use a different model Yes, this is a very good thing when "a different model" means "a worse model." > People who want to deceive or manipulate are not using Google models anyways. They are going to use a model without safety rails That's totally invalid logic. There are plenty of deception and manipulation use cases that don't run afoul of model safety rails at all. Trivially: Creating fake dating profiles to scam people. Fake product images. Fake insurance claims. Fake blackmail (e.g. of a person and another man/woman at a bar).
- jacquesm 7mo agoWhat a great thing this didn't exist in the past. We likely wouldn't have had any of the amazing artworks that we have now. Imagine an AI generated Mona Lisa, Nightwatch or Sistine Chapel ceiling because prompting would have been so much cheaper than paying Leonardo, Rembrandt or Michelangelo... Now extrapolate to all other artforms. Sculpture seems safe, for now, but only barely so.
- tom1337 7mo agoI'd say these models only exist because we had amazing artworks in the past.
- jacquesm 7mo agoAbsolutely.
- wordpad 7mo agoI feel like the complete opposite is true. Artists aren't doing it for the money. With advanced tools like these they wouldve iterated much faster and created much grander designs. Art is about pushing limits of what's possible and AI just raises those limits.
- __alexs 7mo agoThere is a tremendous amount of "art" that is produced for purely commercial reasons. It employs many thousands of people. These roles are definitely threatened by image generators. Agree that if you are Artist this is not going to be a big concern to you.
- gm678 7mo agoAlso, many (I would even venture to say most) of the great artists most people know of earned their bread with intermittent commercial contracts, even rote advertising commissions in the 19th/20th century.
- coldtea 7mo ago
- hmokiguess 7mo agoThe Chinese are so much ahead in this space, their models are way better at this stuff. For example, https://hunyuan.tencent.com/image/en?tabIndex=0 https://hunyuan.tencent.com/image/en?tabIndex=0 and https://seed.bytedance.com/en/seedream5_0_lite https://seed.bytedance.com/en/seedream5_0_lite
- KK7NIL 7mo agoKind of a pointless comment without a link to such a model.
- hmokiguess 7mo agosorry, I've edited my original comment with one such example, but they're easily discoverable I assumed it was popular knowledge - https://hunyuan.tencent.com/image/en?tabIndex=0 https://hunyuan.tencent.com/image/en?tabIndex=0 - https://seed.bytedance.com/en/seedream5_0_lite https://seed.bytedance.com/en/seedream5_0_lite someone shared benchmarks that differ my experience tho, so I may be biased
- sigmar 7mo agowhich models? we have user-preference rankings that put NB2 on top: https://arena.ai/leaderboard/text-to-image https://arena.ai/leaderboard/text-to-image
- hmokiguess 7mo agointeresting, maybe this is just anedoctal experience then and I'm biased but I have been preferring theirs over Nano Banana
- raincole 7mo agoWhen it comes to image prompts, Seedream is far behind Nano Banana. "Far behind" is a ridiculous understatement here, btw. Afaik the only real competitor is Riverflow V2.
- evrenesat 7mo agoI only needed help of this banana boy twice, it managed to disappoint me each time. The most recent one, I was trying different beard and mustache styles on myself, on a photo I imported from my own Google photo gallery, and it consistently rejected me, claiming I'm a public figure. Nobody ever told me that I look like any famous person, so that's googles own bananination. ChatGPT nicely handled the job.
- wnevets 7mo agodoes it still break images with transparent pixels?
- fasteddie31003 7mo agoI'm building my personal home right now. The AI image models have been a game-changer in designing the look of the house. My architect did an OK job, but the details that Nano Banana added really bring the house up a notch. I just do hundreds of renders from the basic 3D models and I find looks that I like and iterate from there. We are implementing the renders from Nano Banana over our Interior Designers designs. We would not have hired the Interior Designers again after using Nano Banana to do our interiors. I think part of the issue with architects and designers today is that they use CAD too much. It's easy to design boxes and basic roof lines in CAD. It's harder to put in curves and more craftsman features. Nano Banana's renders have more organic design features IMO. Our house is looking great and we're very happy how it's going so far with a lot of the thanks to Nano Banana.
- kristjansson 7mo agoPart of the job of interior design is delivering the promised images in … yknow, physical reality? How are you going from nano banana images to actual plans, materials, finishes, products, paint codes, … ?
- yokoprime 7mo agoThe interior designer doesn't really do squat. They can do plan drawings and have some off the shelf cupboards and furniture. They don't implement anything
- fasteddie31003 7mo agoI just gave the renders to the cabinet makers and they had no problems recreating.
- kristjansson 7mo agoInteresting. I model interior architecture as "here's $xxxK, make it nice" and they do a bunch of work to figure out what you mean by nice, and a bunch more work to codify your definition of nice into, like, SKUs of sconces and so on. Seems like NB helped you figure out your definition of nice, and your subcontractor had a good designer on staff to execute on that.
- throwaway4928ab 7mo agoCan we now edit the images it spits out? All prior tests in trying to edit AI images has failed miserably and laughably
- vunderba 7mo agoI've only had a brief opportunity to try out NB Pro 2 (`gemini-3.1-flash-image-preview`), so I haven't had a chance to update GenAI Showdown. Here's some of my captions that tend to trip up even state-of-the-art models. https://mordenstar.com/other/nb-pro-2-tests https://mordenstar.com/other/nb-pro-2-tests So far it does feel more iterative than an entirely new leap in terms of capabilities, but I haven't run it through the more multimodal aspects such as editing existing images. That being said, it actually managed the King Louie jump rope test which surprised me.
- UnknownBanana 7mo agoI love your website, art and projects!
- vunderba 7mo agoThanks I appreciate it!
- dialogbox 7mo agoNice test. Nitpicking. Isn't it NB Pro and NB 2? Not NB Pro 2.
- vunderba 7mo agoYou're totally right. THROW THE WHOLE THING OUT! Seriously thanks - will update the site.
- runamuck 7mo agoI saw an item for sale on Ali Express's video and I thought "Wow, they hired some really attractive actors to pitch their little gadget." 30 seconds in, I realized they used GenAI. Not because it looked AI, but because the production values looked too high and professional for the item. I would get in on this if you sell anything online.
- coffeebeqn 7mo agoThey can even combine the models, create the presenters with nano banana and then use that as the reference for a video model and paste in your product
- arctic-true 7mo agoOne thing I notice is that the voices in video AI are absolute hogwash. Voice AI is great, video AI is great, but AI videos where humans speak give me the feel of really poorly dubbed foreign TV - the timing is not quite right and the facial expressions don’t always match up with the words being spoken.
- ge96 7mo agoMy naive question, can image generation make something novel eg. "show me a DNA structure that cures cancer" can it do that, or it has to have seen something before to generate it. Just think we conceptually know what a brushless motor design looks like and it's just pixels. I guess even if it did produce the image we wouldn't know what it means.
- minimaxir 7mo agoAll image models can generate images that were not in its training dataset, but it can't generate reductive extreme cases like your example.
- ge96 7mo agoWhat about it is extreme? It's a concept, like "generate an xray image" eventually hopefully the cure to cancer could be represented as a simple molecule or whatever, I'm not saying I know.
- minimaxir 7mo agoThere is currently no knowledge nor progress for what a cure for cancer, and nothing a LLM can draw upon. You could generate "pregnant Elon Musk with four arms and three eyes doing yoga poses" because the image models have enough visual concepts of each of those individual things, but that specific image is (likely) not in any training dataset.
- ge96 7mo agoWhat I'm saying is if this thing can generate random things (noise) couldn't it make that or new tech like negative mass. Anyway I get it too if we don't know then something we made wouldn't know.
- claysmithr 7mo agoYou are overestimating it's intelligence, but I bet it would hallicinate some result, why not try it yourself?
- CWuestefeld 7mo agoWhat they've chosen as examples to illustrate the strength of the new model surprises me. The "cubism" example seems like it would be a closer fit to something like stained glass or something. I don't think the thing really understands what cubism was all about. Cubist painters were trying to free themselves from the confines of a single integral plane of perspective by allowing themselves to show various parts of the image from different viewpoints, different times, different styles, etc. The division of the image into geometric shapes is just a by-product of that quest, whereas the examples here have made it the sum total of the whole piece. This feels to me like an example of how LLMs still don't "understand" what the art means, and are just aping its facade.
- kevinsync 7mo agoI had a similar thought before realizing that I'm pretty sure what they were demonstrating wasn't art style, but adherence to correct physical dimensions and construction of the buildings referenced, that was then expressed in an art style (or reasonable facsimile thereof). The before prompts would just conjure a random building out of thin air, the after prompts searched the web for reference material and then used that in image generation. And actually, the link I saw a bit ago was this [0] which is more in-depth and has a lot more examples + prompts. [0] - https://deepmind.google/models/gemini-image/flash/ https://deepmind.google/models/gemini-image/flash/
- zug_zug 7mo agoI'm sure this has been written about but here's what happens long term - images are commoditized and lose their emotional appeal. Probably about half of us here remember photos before the cell phone era. They were rare, and special, and you'd have a few photos per YEAR to look back on. The feel of photos back then, was at least 100x stronger than now. They were a special item, could be given as a gift. But once they became freely available that same amount of emotion is now split across many thousands of photos. (not saying this is good or bad, just increased supply reducing value of each item) With image/art generation the same thing will happen and I can already feel it happening. Things that used to be beautiful or fantastic looking now just feel flat and AI-ish. If claymation scenes can be generated in 1s, and I see a million claymation diagrams a year, then claymation will lose its charm. If I see a million fake Tom Cruise videos, then it oversaturates my desire for desire for all Tom Cruise movies. What a time to be alive.
- nathan_compton 7mo agoPeople here like to say "Commoditize your Compliment" but to a company the size of google or amazon literally EVERYTHING is your compliment. Too bad no philosopher or political scientist or economist every thought about this stuff before or we might have some kind of plan to make the future less miserable and alienating.
- deleted 7mo ago[deleted]
- NoGravitas 7mo ago> Too bad no philosopher or political scientist or economist every thought about this stuff before I see what you did there and know exactly the political economist you are talking about, but if you Speak His Name, the shrieking hordes descend.
- deleted 7mo ago[deleted]
- GaggiX 7mo agoYou can still buy a Polaroid, there is one factory left in the world able to produce the film required but they still make them.
- danesparza 7mo agoIs it just me, or is Nano banana not working in Gemini currently?
- jakub_g 7mo agoSince talking images, are there any AI models that can output real transparent gifs/pngs? And not a (botched) fake white/gray grid background that is commonly used to visualize transparency?
- minimaxir 7mo agoYou can output to a plain background and use any number of tools to mask it.
- jakub_g 7mo agoI know. It sounds like a perfect task for AI to do it though (wasn't the whole premise of AI do to mundane things for us), yet they fail to do it, and I need to use an external tool.
- minimaxir 7mo agoAlpha is a 4th image channel that 99%+ of images in the training data do not use, so it makes more pragmatic sense to just not allow it.
- Zarathruster 7mo agoWeirdly, MSPaint (or whatever it is now) is really good at this (with, I assume, an AI model)
- deathanatos 7mo agoThe output from Nano Banana, even when it is ostensibly drawing "single color" shaded areas, is so jittery that it can be challenging to threshold it.
- minimaxir 7mo agoNano Banana Pro is better at that type of thing, so I suspect Nano Banana 2 will be sufficient.
- 7mo ago
- LeoPanthera 7mo agoIt's notable that this model is less advanced that the previous "Pro" model, and also that the Gemini interface is defaulting all requests to "Fast" even if you've previously changed it to Pro. I guess even Google is running out of GPUs.
- JOJESU 7mo agoI’ve been exploring this exact problem space from the angle of extreme constraints (single-digit MB memory, no cloud assumptions). I documented what broke first and why here, in case it’s useful: https://github.com/nullclaw/nullclaw https://github.com/nullclaw/nullclaw
- nathan_compton 7mo agoSo this is an ultra-minimalist software platform to farm work out to enormous energy chugging AI models?
- RivieraKid 7mo agoIt's extremely slow, takes several minutes to generate an image.
- hedora 7mo agoOpen weight? How many parameters?
- Scene_Cast2 7mo agoIt still seems to have the same pitfalls as all the other image generation models. I ran it through my test prompt (wary of posting it here, lest it gets trained on) - it still cannot generate something along the lines of "object A, but with feature X from Y", where that combo has never been seen in the training data. I wonder how the "astronaut riding unicorn on the moon" was solved... EDIT: after significant prompting, it actually solved it. I think it's the first one to do so in my testing.
- jslakro 7mo agoIt'll be great to find a web directory dedicated exclusively to good/useful prompts with nano banana
- hubraumhugo 7mo agoIt's working pretty well for generating an xkcd comic for your HN profile: https://hn-wrapped.kadoa.com/ https://hn-wrapped.kadoa.com/ Previous nano banana frequently made speech attribution errors, the new one seems a lot more consistent.
- lightyrs 7mo agoThis was really fun. Nice job.
- neom 7mo agoI did some tests, my education is in digital imaging technology/film from 20 years ago so I find this stuff fun to follow. Two what I could consider "interesting prompts" for image gen testing. Did pretty well. https://s.h4x.club/eDuOzPDd https://s.h4x.club/eDuOzPDd "A macro close-up photograph of an old watchmaker's hands carefully replacing a tiny gear inside a vintage pocket watch. The watch mechanism is partially submerged in a shallow dish of clear water, causing visible refraction and light caustics across the brass gears. A single drop of water is falling from a pair of steel tweezers, captured mid splash on the water's surface. Reflect the watchmaker's face, slightly distorted, in the curved glass of the watch face. Sharp focus throughout, natural window lighting from the left, shot on 100mm macro lens." - Only major problem i could find at a glance is the clasps don't make sense probably, and the drop of water inside the watch on the cog doesn't make sense/cog mangled into tweezers. https://s.h4x.club/yAuNPlRk https://s.h4x.club/yAuNPlRk "A candid photograph taken from behind an elderly woman sitting alone on a park bench in late autumn. She is gently resting one hand on the empty seat beside her, where a man's weathered flat cap and a folded newspaper sit untouched. Fallen golden leaves cover the path ahead. The low afternoon sun casts her long shadow alongside a second, fainter shadow that almost seems to be there, the suggestion of someone sitting next to her, visible only in the light on the ground. Muted, warm color palette, shallow depth of field on the background trees, photojournalistic style." - I don't know why but it internal errored twice on this one but then got there.
- ozgung 7mo agoAny info or speculation about technical details?
- keiferski 7mo agoSome random predictions about what AI image generation tools will do/are doing to art: 1. The narrative/life of the artist becomes a lot more important. The most successful artists are ones that craft a story around their life and art, and don't just create stuff and stop. This will become even more important. 2. Originality matters more than ever. By design, these tools can only copy and mix things that already exist. But they aren't alive, they don't live in the world and have experiences, and they can't create something truly new. 3. Those that bother to learn the actual art skills, and not merely prompting, will increasingly be miles ahead of everyone else. People are lazy, and bothering to put in the time to actually learn stuff will stand out more and more. (Ditto for writing essays and other writing people are doing with AI.) 4. Taste continues to be the single most important thing. The vast, vast majority of AI art out there is...not very good. It's not going to get better, because the lack of taste isn't a technical problem. 5. Art with physical materials will become increasingly popular. That is, stuff that can't be digitized very well: sculpture, installation art, etc. Above all, AI art is uncool, which means it has no real future as a leading art form. This uncoolness will push people away from the screen and towards things that are more material.
- avmich 7mo agoI mostly disagree. > 1... The narrative/life of the artist becomes a lot more important. When I watch a movie, I don't care about the artist's life. I care about character life, that's very different. > 2... Originality matters more than ever. By design, these tools can only copy and mix things that already exist. It's like you assigning to humans divine capabilities :) . Hyperbolizing a little, humans also only copy and mix - where do you think originality comes from? Granted, AI isn't at the level of humans yet, but they improve here. > 4... It's not going to get better, because the lack of taste isn't a technical problem. Engineers are in business of converting non-technical problems into technical ones. Just like AI now is way more capable than it was 20 years ago, and able to write interesting texts and make interesting pictures - something which at the time wasn't considered a technical problem - with time what we perceive as "taste" may likely improve. > 5... Above all, AI art is uncool, which means it has no real future as a leading art form. AI critics are for a long time mistaking the level with trend. Or, giving a comparison with SpaceX achievements, "you're currently here" - when there was a list of "first, get to the orbit, then we'll talk", "first, start regular payload deliveries to orbit, then we'll talk", "first, land the stage... send crewed capsule... do that in numbers..." and then, currently "first, send the Starship to orbit". "You're currently here" is the always existing point which isn't achieved at the moment and which gives to critics something to point to and mount the objection to the process as a whole, because, see, this particular thing isn't achieved yet. You assume AI won't be able to make cool art with time. AI critics were shown time and time again to be underestimating the possibilities. Some people find it hard to learn in some particular topics.
- dyauspitr 7mo agoI really wish they opened a version of this up for adult content. They would make immense amounts of money and it could be fenced off behind some sort of paywall where they could verify the age of the person.
- MaxikCZ 7mo agoI have Google AI Ultra. Where can I test this? They say its in aistudio, which says its a paid model and I need to setup billing (as if paying for Ultra isnt enough). They say its available in antigravity, but I cant seem to find it there?
- Sevii 7mo agoWorks for me with a Pro sub at https://gemini.google.com/app https://gemini.google.com/app
- rosstex 7mo agoAdding to predictions: the magic of travel might actually be reborn, as people seek authentic experiences.
- mowmiatlas 7mo agoSo, I'd suspect the seedance2.0 competitive video model is coming as well soon? ;)
- monster_truck 7mo agoKind of surprised it hasn't been pulled yet. Have seen some very disturbing (grok tier) examples of completely bypassing whatever censors they have in place by simply asking gemini to write the prompt
- Invictus0 7mo agoIt's not working very well at all. I started with a picture of a girl sitting at a cafe table and asked it to zoom in, and it enlarged her head to the size of a balloon.
- userbinator 7mo agoLOL! It either has developed a sense of humour, or your prompt was not specific enough.
- casey2 7mo agoStill has context leaking into the text/random signs in the image, made worse by generating filler with an internal LLM
- thinkingemote 7mo agoIf any AI image generation companies are reading this, I want the image to be in layers which can also be exported, so I can 1) do post processing of my own or 2) arrange for an AI image generation model to process just the layers i specify.
- grallm 7mo agoYou can use Qwen Image Layered to split the image into layers https://qwen.ai/blog?id=qwen-image-layered https://qwen.ai/blog?id=qwen-image-layered
- tariky 7mo agoThis looks like a response to Seedream 5.0 lite that was published two days ago. I use all those fancy image models editing capabilities for my fast fashion web shop. I must say: product photography for clothing and accessories product is dead. Those models are amazing at style transfering and garment transferring. We will see how good will be Seedream 5.0 full version.
- Tiberium 7mo agoSeedream 5 Lite is honestly extremely disappointing, its text to image is way worse than 4.5, image editing is fine but that's it. It's way, wayy behind NB2.
- mattlondon 7mo agoYou think they created and launched this in just 2 days? These things take a lot of effort and time to develop...
- zhyder 7mo agoModel card: https://deepmind.google/models/model-cards/gemini-3-1-flash-image/ https://deepmind.google/models/model-cards/gemini-3-1-flash-... Pretty close to Gemini 3 Pro Image (aka Nano Banana Pro) in most benchmarks, even without thinking+search, and even exceeding it in 2 most important ones of 'Overall Preference' and 'Visual Quality'. I'm excited about the big jump in Infographics/Factuality (even without thinking+search; I'm surprised that text+image search grounding doesn't make an even bigger dent).
- CrzyLngPwd 7mo agoJust what we need, more sloperators thinking they are being creative and making art by prompting. I would be happy to never see any more AI slop.
- jorvi 7mo agoThis will stay useless for editing personal pictures so long as virtually every prompt with a person in it is met with "I can't edit images of some people". For whatever reason, they've made the celeb detection so ultra-aggressive that almost everyone is detected as a (lookalike) celeb.
- Tiberium 7mo agoIt's only for Europe, you should try a US VPN or, in the worst case, use it over Vertex AI, which allows you to generate anyone.
- geooff_ 7mo agoFunny timing. I just migrated my personal styling app off of Nano Banana. My main use case is editing user uploads to enhance their clothing images. A large part of it is preserving logo, graphics and other technical details. I noticed over time it felt like Nano Banana has gotten worse at this. I have a test set of graphic t-shirts that I noticed the model seeming getting worse with it. This combined with price and the terrible experience of their cloud console got me to migrate off.
- tiffanyh 7mo agoKeeping track of the different AI product names is so confusing even from a single company. Why can't Google, for example just call: Gemini Image = Nano Banana Gemini Video = Veo ...
- hadley 7mo agoLet alone that Nano Banana 2 is Gemini Image 3.1
- h4ch1 7mo agoWow it's capable of critiquing its' own output? <OUTPUT> While the overall aesthetic matches the minimal white-stroke style and technical design you requested, and the provided step descriptions are included, please note that there are a few minor rendering artifacts in this specific generation: The text on the banner entering the vault in step 8 is illegible. There is a small typo in the caption for step 6 ("CONFLSCT" instead of "CONFLICT"). Despite these small imperfections, this layout should work well as a guide for your canvas implementation. </OUTPUT>
- divan 7mo agoWhat is annoying about Nano Banana, is how bad is experience when you try to iterate or, especially, repeat same task for another photo. After second of third image it starts randomly ansering with complete nonsense like "I'm just a language mode and can't assist with that" or "I can't do that" (with absolutely the same prompt it had no issues 2 photos in a row in the same chat). It also gaslights me, when I point out on an error. I tried to create a cartoon portrait of the person from photo and use background from another photo. It got wrong the order of photos. I provided filenames and explicitly told which one is for person and which for bg. It generated it wrong again, and all attempts to explain that it got it wrong were met with "No, it's YOU incorrect". So frustrating.
- Zarathruster 7mo agoThis reminds me of my experience trying to generate a reference photo for a 3d model. I told Nano Banana to generate an image of the character with his feet shoulder width apart. It ended up generating him with his feet pressed together, so I told Nano Banana to widen his stance slightly. It gave me an image of the man with his feet spread far apart enough to straddle a horse. I asked for a slightly narrowed stance and his feet were once again brought together. This went back and forth unsuccessfully for a while until I asked, "I'm asking you to make his feet shoulder-width apart. Why are you ignoring me?" And Nano Banana confidently asserted that they are shoulder width apart, and I must be wrong. Ultimately I ended up telling the model to render the same character, pinching a cantaloupe between his ankles, and then to remove the cantaloupe. It worked, but why do I have to trick Google's SOTA image generator to give me very basic stuff like this?
- deafpolygon 7mo agoThe prompt “Generate an image of an orange man riding a bicycle.” no longer generates Trump riding a bicycle. I’m guessing they fixed that. But the prompt "can you depict a cartoonish orange man with a pooh bear in political cartoon style?” correctly generates Trump.[1] So there’s that. [1]: https://imgur.com/a/Gvm5Zje https://imgur.com/a/Gvm5Zje
- vunderba 7mo agoResults are in for `gemini-3.1-flash-image-preview` (NB 2) for the GenAI Showdown site in the editing comparisons. Remember to click the "Pass/Fail" button to toggle between pass/fail and a weighted score to account for additional factors like steerability, image quality, etc. Unfortunately, unlike the leap from NB to NB Pro, we did not see significant gains from NB Pro to NB Pro 2. In several cases (such as the Jaws Poster), we observed that it was substantially more difficult to prevent NB Pro 2 from making significant changes to the rest of the image. Localization of edits, in general, seems to have changed and not necessarily for the better. http://genai-showdown.specr.net/image-editing http://genai-showdown.specr.net/image-editing Comparison solely between the Gemini models (NB, NB Pro, and NB Pro 2): http://genai-showdown.specr.net/image-editing?models=nb,nbp,nbp3 http://genai-showdown.specr.net/image-editing?models=nb,nbp,...
- notnullorvoid 7mo agoThere's something odd feeling about many of the latest AI image models (NB and NB2 included), they feel almost too polished, calculated, and rigid. It doesn't take much effort to get some art with a sense of originality or flow from Stable Diffusion, but I can't seem to prompt NB to get similar results. SD definitely struggles with details especially when trying photo realism, but it'd be nice if newer models could get back some of the roughness that made SD good.
- Zarathruster 7mo agoThis has been my experience too. The new models don't spit out nightmare fuel as often, but they aren't nearly as creative either. They seem to be very good at creating stock photos and not much else
- hwj 7mo agoIs it possible to run this locally on Apple aarch64? (Sorry, I'm probably one of the few HN users left that don't have much experience with AI).
- hattimaTim 7mo agoGood but the gemini api is unreliable as hell. Why would you give a paid user "Resource Exhausted " errors when you have enough resources for free users?
- antkim 7mo agocreate a color pencil style artwork that is extremely detailed, with a surrealistic style of a farmhouse near a cornfield. There should be a cow or two, some chickens and a pig pen with piglets scurrying around. The sky should be overcast and threatening a storm. Purple mountains can be seen far in the distance
- antkim 7mo agoCreate a drawing that looks like a detailed color pencil drawing. It should be of a farmhouse near a cornfield. There should be a couple of cows, some chickens and a pig pen with piglets scurrying around. The sky should be overcast and threatening a storm. Some purple mountains can be seen in the distance.
- nektro 7mo agogenuinely one of the most evil things humans have produced. shameful.