7 ms·
Qwen Image 2.1
- weee322 15d ago[flagged]
- fishfasell 15d agoThe capabilities of local LLM text-to-image is honestly pretty damn impressive. IMO, I think local image generation is currently ahead of local code generation. I can get an image in seconds locally with the quality being way higher than what I'd expect from a local model. However with coding it's much slower and much less impressive. I'm sure there's a reason for this and I'm not an AI expert so I'll let the smarter folks tell me why, but that's just been my observation thus far.
- victorbjorklund 15d agoI mean I’m sure it’s the reverse for an artist. They would be less impressed with the image and more impressed with the code quality
- gedy 15d agoTo generalize, LLMs are great at what you are not skilled at.
- fishfasell 15d agoThat's a fair statement, I agree. I'm quite an abysmal artist so I could be a victim of my own bias here
- 26d0 15d agoThe point I think is interesting is that this is just 7B. The current SOTA 7B LLMs are barely usable for quite simple coding.
- becquerel 14d agoText is in a sense way harder to do than images because of radical nonlocality. A word at the start of one paragraph can directly influence the meaning of a word five paragraphs away. Whereas images typically represent the real world, or at least a spatial domain, which gives you a lot of structure 'for free'. If you are drawing a human, you can make a reasonable guess where their hands go in relation to their face. If someone hands you the first half of an essay, finishing it is not trivial.
- mft_ 15d agoI've played with diffusion models on and off since the first release of Stable Diffusion - just for amusement, without a particular goal. Recently, I've been helping a friend's wife with some basic vector images for her sewing hobby (she has what is essentially a CNC sewing machine) and have been super-impressed with FLUX.1-Kontext, which I've been running on my Macbook Pro with mflux. Its ability to (for example) take a photo of a human or an animal and return a line drawing which is recognisably them (rather than just a generic similarish image as I've experienced with other models) is excellent. It's an older model now, but (AIUI) has the text-handling features baked in, and in my various testing is very reliable at giving me the outputs that I want, without the randomness I've experienced previously. It's big and relatively slow (~3 mins per 512x512 image edit on my M1 Max Mac) but excellent to work with. It's also very straightforward to set up, without the harness complexity of e.g. comfyui.
- jLaForest 15d agois the cnc sewing machine an off the shelf model or something DIY? I'd love to hear more
- mft_ 15d agoOff the shelf - it’s a Brother. It prints via a proprietary file format (.PES) but there’s an extension for Inkscape that supports creation and export.
- agentdev001 14d agoSounds ripe for vibe... sewing
- alirezaxdehghan 14d agoI think they meant an embroidery machine
- rahimnathwani 14d agoHow are you converting the bitmaps into vector images?
- gavmor 15d agoRemember that quality output is a necessary but insufficient property of a generative model. Prompt-adherence is really hit-or-miss—especially if one lacks the visual vocabulary. Likewise with coding, I find junior devs don't think to prompt re: respecting this-or-that interface, or refactoring to point-free style, etc. So, as others have said, the artist knows better.
- tarcon 14d agoI think there was a lot more brainpower invested in the media generation side of things. The noise-based diffusion technique is further developed. It had a discovery of applying a physics-based understanding of Brownian motion to guide it. Image generation has comparatively simple training process - this is an image with dog, and without dog (contrastive learning). Might be worth to watch the diffusion based LLMs.
- hn45e7pbij 15d agoImage gen you eyeball one frame and stop, code needs hundreds of tokens all correct in sequence, one bad line and the whole thing fails.
- jacktu 14d ago[dead]
- spottedmarley 15d agoBoy do I love waking up to find a new awesome toy from the Qwen team waiting for me to play with! Pulling it now
- Hard_Space 15d agoInteresting in the example of assembling the Cheers team how the otherwise great result genericizes Shelley Long.
- hughc 15d agoThe result seems a pretty good representation given the source image wasn't that great. I think that Woody Harrelson comes across much worse.
- TomGarden 15d agoVery impressive, and kind of worrying a 7B model can have such capabilities. The implications are huge. And Qwen does no watermarking (yet) yeah?
- d2kx 15d agoGod I love the Qwen team. Easily the most diverse set of models from all the Chinese labs. Only Gemini/DeepMind comes close.
- mdp2021 15d agoHow do you use this model locally, similarly to using `llama-server -m <model>`? (I mean: outside direct or substantial use of Python, and running the Neural Network in the most efficient way.)
- embedding-shape 15d agoProbably ComfyUI is one of the easiest way to get started with local image/video models. Or perhaps vLLM, if they have support for it already, would be something like `vllm serve <model> --omni --port 9080`
- utopiah 15d agowhy not just as you suggested i.e. https://qwen.readthedocs.io/en/latest/run_locally/llama.cpp.html#llama-server https://qwen.readthedocs.io/en/latest/run_locally/llama.cpp.... then get the result either via a UI or wget/curl it back?
- mdp2021 15d agoI am not sure that llama.cpp also supports image generation models.
- utopiah 15d agoit's multimodal, see https://github.com/ggml-org/llama.cpp/blob/master/docs/multimodal.md https://github.com/ggml-org/llama.cpp/blob/master/docs/multi...
- exe34 15d agoMultimodal doesn't guarantee input and output. > Currently, we support image, audio and video input.
- utopiah 15d agoSeems I'm missing something. Does this model support other inputs? Image outputs are supported, videos I'm not sure but I don't think that's an output, just a preview of the equirectangular example, so, same question here, what does this model outputs that isn't supported?
- trains39472 15d agoA 7B diffusion model can now render CJK text better than Microsoft Windows.
- tomjen3 15d agoJust think about how recently we got that feature in the official ChatGPT image gen. And now we have that running locally — assuming that is, I can figure out how to get this running on my Mac — blows my mind.
- Havoc 15d agoPretty sure comfyui has a mac executable
- a96 14d agoFSVO executable, yes. https://formulae.brew.sh/cask/comfy https://formulae.brew.sh/cask/comfy
- jimmydoe 15d agoIs ChatGPT really that good? Back in Apr, ChatGPT Images 2.0 has some broken Chinese texts in its featured examples, and they later removed that from blog post. Is 2.5 better now?
- tomjen3 15d agoI want to say that they released a version since then before 2.5, but I'm not entirely sure. I should also note that I do not know a single Chinese character, so it's possible they are broken in such a way that I wouldn't necessarily notice. I do remember, however, that English text used to be broken; now it's completely readable.
- doctorpangloss 15d agoIdeogram 4 has been around for a while haha
- hgufj 15d agoI am really grateful to the Chinese Labs for open sourcing their best models. If it was left to the Americans, we would be forced to pay obscene API fees to use them.
- Gasp0de 14d agoWhich lab is open sourcing their model?
- jfoster 15d agoNote that the license on this has this in it: > You shall not use the Materials for any commercial purpose without obtaining a separate commercial license from us. It probably will be much cheaper to use than other image models, but it seems that will be up to the whims of Qwen/Alibaba rather than just being the cost of putting it in a cloud provider. https://github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE https://github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE
- tenuousemphasis 15d agoGood luck to them enforcing that license.
- RIMR 14d agoHonestly, that's fine. The commercial license isn't that bad, and cloud providers selling API access to this can afford it. I am just happy I can run these models on my own hardware. Hopefully in 10 years, self-hosted models far exceeding what's currently available will run comfortable on commodity hardware.
- jfoster 15d agoA lot of the previous Qwen models seem to have used Apache licenses, among others: https://en.wikipedia.org/wiki/Qwen#List_of_models https://en.wikipedia.org/wiki/Qwen#List_of_models Unfortunately, it looks like this model is using a much more restrictive license: https://github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE https://github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE
- gregoriol 15d agoWas going to post about this: the last image models with Apache 2.0 license seem to be from 2025, recent Qwen models are "non-commercial use".
- bloaf 15d agoI love the non-commercial clauses because of how many people are using these for deceptive ads and “virtual staging” and fake social media accounts. Anything that makes those guys lives harder while still letting me make silly pictures for my kids and tapestries for my D&D campaign feel fine by me.
- tenuousemphasis 15d agoYou think they care about the probably unenforceable license terms?
- user43928 15d agoThis achieves absolutely nothing to that end. People can continue to use closed SOTA models to generate outputs for commercial or malicious purposes. What this research license achieves is that we cannot use this model in applications we publish.
- Luker88 15d agoCompanies can use llm to license-wash open source code regardless of license. How difficult would it be to use this model to create a second model without licensing issues?
- deleted 15d ago[deleted]
- trentor 15d agoThey finally fixed their VAE. It really held back their models over the last 2 years. EDIT: It still produces artifacts it's better but unusable for production work. In midvalues you will see a slight dot pattern.
- mdp2021 15d ago> finally fixed their VAE Can you share the sources?
- trentor 15d agoIt's right there in the hugging face link? latents go from 16ch @ 8x compression to 64ch @ 16x, so roughly the same total latent budget but much more channel heavy. It’s also deeper/wider, and the old 2x2 transformer patching is gone. On some images it still produces artifacts but can't say if it's the transformer or the VAE yet.
- Ristovski 15d ago> In midvalues you will see a slight dot pattern. Is this not simply some sort of watermark instead of an artifact?
- trentor 15d agoNo, it's probably their rope implementation. They had a similar problem with the old qwen image but to be sure it needs some digging.
- BlackGlory 15d ago[flagged]
- gunalx 15d agoIts happy to see a new open image model from qwen. But the license is a let down. And it dosent even beat their closed qwen3 image wich is already a bit old.
- yorwba 15d agoQwen Image 3 was released two months ago: https://qwen.ai/blog?id=qwen-image-3.0 https://qwen.ai/blog?id=qwen-image-3.0 I think you have it confused with another model.
- bknight1983 15d agoWhile I'm impressed with the Bluey example, the lack of Muffin disappoints me.
- docheinestages 15d agoMy first impression is that it's not so good at following prompt directions. I asked it to place a 3D text made of glass in a particular city. It instead gave me a broken 3D text on a white background. Maybe with different seeds it gets better, but it's more of a trial and error process than reliable results.
- dannyw 15d agoTry translating your prompt to Chinese first, it seems a lot better at understanding and following Chinese prompts even with the translation hop.
- mdp2021 14d agoYou could try attaching other images as references (I think you can attach a maximum of 10 images). If the attachments can be blurred or sketchy or generic enough, they could be used for generalization.
- samayashar 15d agoQwen and Alibaba are the biggest competitor for basically every model out there. They're beating the benchmarks like top-frontier models, focused on open-source and much cheaper than the competitors. Excited to see what the future holds for them!
- jjcm 15d agoI run a prompt-to-ui design site that uses image models for the design process[1]. The text rendering especially makes this model deeply interesting to me, despite the license. Here are some tests using my harness comparing the outputs of gpt-image-2 and qwen 2.1: https://html.non.io/qwen-comparison/ https://html.non.io/qwen-comparison/ The text rendering definitely is much, much better than anything else on the open weights market right now. Small text fidelity is quite good. It seems like the text encoder however gets a little bit overloaded with larger prompts - note the presence of hex codes in the design output, those were inputs from the expanded prompt. I'll be trying a post-training run on this for web design, it has some serious potential. [1] diffui.ai
- xienze 15d ago> The text rendering definitely is much, much better than anything else on the open weights market right now Really? Because basically everything in those screenshots is completely garbled. I didn't follow it super closely but I thought Ideogram or whatever was really good for this particular use, with actual clear text.
- vunderba 15d agoThis is my experience as well. Ideogram4 (assuming you are willing to put in the work to use the proper structured JSON input) is very accurate when it comes to text rendering in an image.
- cloudking 15d agoThose simple prompts produce nearly the exact same layout in the 2 different models?
- colesantiago 15d agoWhile the license of this model is a shame it is still unenforceable. I know a few friends of mine who are running models and are ignoring the licence. Whether it is AGPL 3.0, or a completely restrictive license, it is going to get broken anyway and be used for commercial purposes. I don't know anyone who looks at the licenses of the OSS software they are using. In today’s world OSS is synonymous with "Free" and the AI model providers are proof of that with their training of code, datasets, etc. So it begs the question, why should we abide by their licenses of their models?
- dawnerd 15d agoAnd they shouldn’t be enforceable considering how the training data was slurped up without concern for licensing.
- deleted 15d ago[deleted]
- vunderba 15d agoSo thoughts Positives • It's a heck of a lot smaller than Qwen-Image 1 (20b parameters) at only 7b, making it one of the smaller open-weight models available (Z-Image Turbo is one of the few that is smaller at 6b) when compared to Ideogram, Krea2, Flux2, etc. • It supports native transparency (Qwen's team, as far as I know, is the only one attempting to tackle this). Even though it's relatively trivial to set up background removal postprocessors, it's also neat to see it natively supported. • It's fast using QwenImage2.1 convrot, a 1MP image took around ~5 seconds on an RTX4090. Negatives • The license (assuming you respect it) is far more restrictive. The original Qwen Image 1 was released under the standard Apache license; this one explicitly forbids commercial usage without obtaining a separate license. On the other hand, a lot of us didn't expect the Qwen team to ever release "weights-available" ever again. Qwen-Image 1.0, released about a year ago, only scored 4/15 on my GenAI Showdown Benchmarks. Since that time, they've been upstaged by Krea 2 (6/15) and Ideogram4 (8/15). I'll post the new results once I have some more time to run them. https://genai-showdown.specr.net https://genai-showdown.specr.net
- alightsoul 15d agoThey're trying to cash in but this is just sad
- thenthenthen 15d agoWithin a few days this seems a total pivot from Xiaomi’s op RL dashboard and the praise of Chinese open model? What is the sentiment now?
- Foobar8568 15d agoQwen can't train anymore with openai reasoning tokens? I kid I kid.
- cmrdporcupine 15d agoEverything is combined and uneven, including the opinions of hackernews commenters? There is no single opinion, and clearly no single Chinese approach. Also Chinese labs are in particular very careful about anything which be used to create pornographic content, which is highly illegal in the PRC.
- golutyagi9710 15d ago[dead]
- amelius 15d agoI don't want only cherry picked examples. Show me failure modes too.
- run-good-code 15d ago[dead]
- hirako2000 15d agoWoody Harrelson?
- guideaitools 15d ago[flagged]
- ramesh31 15d agoStill fails to generate smoothly animated sprites, although the native RGBA transparency is nice. Anyone found one that can?
- jonplackett 15d agoIs it just me or are Alibaba / qwen’s websites often appear broken / very slow?
- 1saadcodes 15d ago[dead]
- andrewdb 15d ago[flagged]
- timmytokyo 15d agoNot sure there's a better avatar for the absurdity of AI slop imagery than the "cowboy on horseback". That's a pony with a child's saddle on it, and they've composited a grown man on top of it.
- finnjohnsen2 15d agoI would call this a license trap: Qwen RESEARCH LICENSE AGREEMENT Code on github, models on huggingface, nice intro text: "We are excited to open-source Qwen-Image-2.1 [...]". meh...
- thenipper 14d agoThe uncanny valley cheers is really freaking me out.
- rickreynoldssf 14d agoPeople it renders look Asian. If you give it a reference image of a caucasian it will render an Asian. I wonder why?
- solarkraft 14d agoSame reason photography was historically badly calibrated for black skin. (https://www.shutterstock.com/blog/shirley-card-racial-photographic-bias https://www.shutterstock.com/blog/shirley-card-racial-photog...).
- ouch-blurred52 14d ago[flagged]
- andsoitis 14d agoAre all the humans in those photos fake?
- AbstractH24 14d agoThe minute you start looking at other non-frontier models, you understand why Dario, Altman, and Musk are coming together to say we need regulation. None are as good yet, but what everyone said is coming true - models are not moats. And these folks need an exit (even Msuk whose shares are still locked)
- fahrvrgnugen 14d ago[dead]
- Zaraif13 14d agoAnyone get success editing videos, frame by frame, using image editing AI? Which model works well? In my experience, video models generate videos pretty well but are mid at editing. They actually regenerate the entire video along with the edit. So these models being non-deterministic tweak the rest of the video as well, the parts you hoped would be left not edited. It gets exponentially worse when there are humans in the videos, annoying face distortions and for some reason these models just don't understand fingers.
- qsbuilder 14d agoLove that these models are releasing making it easier to get away from these stupid subscriptions.
- deleted 14d ago[deleted]
- s131ph 14d agois it possible to run on 8gb ?
- habajab 14d ago[flagged]
- pan_lid 14d agoCurious how Qwen Image 2.1 performs on less common datasets compared to SDXL. Always good to see more open models.
- mudkipdev 14d ago^ Bot
- globular-toast 14d agoWhat problem is this trying to solve?
- Devin3162 14d ago[dead]
- meherabhossain 14d ago[flagged]
- sgt 14d agoI'm actually looking for a model that is able to create old school pixel art (like 90s style, games like Sierra and so on). But to this date, I haven't seen anything yet. It all reeks of "AI slop". Maybe this is a good thing, I don't know. But if anyone has any tips, I'd appreciate. It's for my own personal use.
- sinan-faizal 14d agoits opensource?
- ldng 14d agoNow if we could get a Qwen-Image-Layered 2 also at 7B ... ^^ (Or does this model has those capabilities natively ?)
- adwinho168 14d ago[dead]