10 ms·
Moebius: 0.2B image inpainting model with 10B-level performance
- NooneAtAll3 3mo agoI don't understand. Is it available somewhere to try or is it just an ad?
- owebmaster 3mo agoYeah it's great but how do I use it? Edit: I think I found it https://huggingface.co/hustvl/Moebius https://huggingface.co/hustvl/Moebius
- K0IN 3mo agowith this size we could have a interaactive web demo.
- james2doyle 3mo agoLike this? https://huggingface.co/spaces/multimodalart/Moebius https://huggingface.co/spaces/multimodalart/Moebius
- IvanK_net 3mo agoWere you able to make it work? It never works in my case.
- ErrorNoBrain 3mo agohttps://simonw.github.io/moebius-web/ https://simonw.github.io/moebius-web/ choose "samlple image" or upload something then mark something with the mouse (and press 'run inpaint') and it'll work a bit and try to hide it, sorta like that "magic eraser" some newer android phones have
- teroshan 3mo agoUnrelated but when I read inpainting and Moebius I was scared it was related and using the art of the great Jean Giraud [0] a.k.a. Moebius https://characterdesignreferences.com/artist-of-the-week-3/moebius https://characterdesignreferences.com/artist-of-the-week-3/m... [0] https://en.wikipedia.org/wiki/Jean_Giraud https://en.wikipedia.org/wiki/Jean_Giraud
- coldtea 3mo agoScared why?
- teroshan 3mo agoScared for the same reason I found last year's 'Ghibli filter' craze upsetting, I would have personally hated to have seen this artist's legacy used for promoting AI image generation.
- TeMPOraL 3mo agoIn case that happened then the rest of the world would probably appreciate the art, and a subset of it, the artist (and even a small subset of ~whole Internet-connected population is a lot of people). Some silver lining, perhaps.
- solid_fuel 3mo ago> In case that happened then the rest of the world would probably appreciate the art What art? We’re talking about generated pictures, aka slop, not art made by a real human. And I don’t know if you’ve been paying attention but people seem to be pretty tired of the slop. I don’t think it would be appreciated nearly as much as you think.
- TeMPOraL 3mo agoThis definition of "slop" doesn't cut reality just quite at the joints. People are tired of marketing. AI generated slop people are annoyed with, is garbage produced for marketing reasons, and it's distinctly noticeable precisely because all the bottom-feeder marketing houses switched to using it. But it's not the AI itself that's the problem here. Slop was here before, but it was made with cheap protein-based image generators. Silicon-based generators are just cheaper.
- N_Lens 3mo agoThe gallery of their samples is pretty impressive!
- epolanski 3mo agoWhat is the current SOTA for impainting? I have a potential project for my e-commerce where I want to allow users to upload images of their house exteriors and impaint awnings.
- vunderba 3mo agoProprietary? Either gpt-image-2 or NB2. I have an example of interior decorating inpainting where I replaced a large floor-to-ceiling window with a mirror, and the result was pretty impressive using NB Pro from nearly a year ago. https://imgpb.com/ZXkiXV https://imgpb.com/ZXkiXV Locally hostable? For my money I'd argue Flux.2 Klein but Qwen-Edit still puts in the work.
- IAmGraydon 3mo agoAs far as I know, gpt-image-2 doesn't even let you define a mask unless you've already run it through one iteration, and once you do define the mask, it just ignores it 90% of the time. It's utterly useless for inpainting. Also, this and other proprietary models are severely limited in their output resolution. I do agree, however, that the Flux2 family is the SoTA at the moment. Running locally via something like Comfy gets incredible results.
- vunderba 3mo agoYeah definitely. You can do workarounds like drawing circles or using highlighters to create pseudo-masks for use with OpenAI or Google models but it’s really just a visual indication more than anything. If you want real precision (especially for complex polygonal masks), or if you’re concerned about image degradation over multiple edit rounds, you'll slam against the limitations of those approaches. Even with SOTA proprietary models, repeatedly editing and re-uploading an image is like making a copy of a copy of a VHS tape: you're gonna see subtle color shifts and quality loss steadily accumulate. At that point, you either need to put in the manual work in something like Photoshop (bringing elements in as layers and masking them properly) or, as you mentioned, use a model or workflow that properly supports masking.
- 3mo ago
- zb3 3mo ago1) What are RAM requirements? 2) If these are reasonable, a WebGPU demo would be great..
- lifthrasiir 3mo agoThe total model size is about 1.2GB (UNet + SDXL VAE included), so probably about ~3GB?
- delis-thumbs-7e 3mo agoThis is the useful AI stuf. There’s so many usecases this makes possible.
- doctorpangloss 3mo agohow many times have you edited a photo you took on your phone in the last 7 days?
- dogomatic 3mo agoPersonally, about 9 times. Would be higher if it was even easier and cheaper
- stusmall 3mo agoI think 3? I feel like that's often enough. Sometimes it's nice to do a quick dumb ass gag on a whim. If I am anything I am a man who loves a dumb ass gag.
- TeMPOraL 3mo agoHalf a dozen at least. (I'm counting only times I used generative editing options in my Galaxy phone - if I were to take your question literally, it would be "at least once every other day", simply due to rotating and cropping.)
- GL26 3mo agoCould this run locally on a smartphone ?
- rasz 3mo agoIt sure has a thing for chins, jaws and removing weight, looksmaxing build in.
- gspr 3mo agoNitpick: in the showcase on that page, under Comparison of Natural Scenes, Moebius should definitely get a "structural confusion" tag for the back of the surfboard. If other models get deducted for truncating the surfboard, then surely the elongation that Moebius does should count too. Also, what's going on behind the in-painted corner of the house? We'd need to see higher resolution pictures, but I'm not convinced that it too shouldn't get a flag. Likewise with the beach just behind the surfboard. Not terrible, but what gets flagged in the competitors is similar.
- james2doyle 3mo agoThere are some demo spaces using this. This one seems the best (paint your own mask) but it failed on all the images I tried: https://huggingface.co/spaces/multimodalart/Moebius https://huggingface.co/spaces/multimodalart/Moebius
- hex4def6 3mo agoI've been playing around, got it to work, although quality was a bit crappy. Still playing around with the settings that get exposed, but you're welcome to look at : https://huggingface.co/spaces/jonatei/MoebiusDemo https://huggingface.co/spaces/jonatei/MoebiusDemo Note that I'm actively messing with it, so it may break for short periods of time :) It's also running on the free CPU, so it's like 80 seconds per image...
- hari1123 3mo agolot of the photo editors on mobiles have this, maybe even some apps?
- michaelfm1211 3mo ago> The core insight of Moebius can be summarized in a single equation: Synergy × (Architecture + Distillation) = Shattering the "Impossible Triangle" of Low Parameters, Fast Inference, and High Quality Is it just me or is it weird seeing these clickbaity AI-generated taglines in an otherwise scientific work?
- dormento 3mo agoIt IS weird, but it "converts" (ugh...), that's why they coming. Apart from this, the text details amazing work. Congrats.
- soperj 3mo agoAfter "In Good Company" i can't hear (or see) the word Synergy without cringing.
- kevin_thibedeau 3mo agoIt signals a paradigm shift in vacuous prose.
- Jackson__ 3mo agoJudging by the performance of the shown examples, the quality is closer to pre-2022 Photoshop content aware fill than actual 10B models. I think it is safe to say this is pretty far from a "scientific" work.
- lifthrasiir 3mo agoTried a bit, and while it is very impressive for 0.2B model it would be very hard to convince me that this matches with 10B models. It did work reasonably well with natural images but inpainted regions were visibly smoother than surroundings, and performed very badly on novel objects. It is also limited to 512x512 output, which limits its practical usefulness.
- amelius 3mo agoDo you think the provided examples are representative of its performance, or do you think they were cherry picked?
- lifthrasiir 3mo agoGiven its limited output dimension it's hard to tell. I haven't exactly tested fine-tuned variants but I think they would work well under certain situations. After all, some (possibly cherry-picked) examples still exhibit similar problems when you inspect them in detail.
- xrd 3mo agoI did an inpainting project for a client a few years ago. They were trying to inpaint banner ads for concert promoters, and find a way to make it easy to produce a bunch of different sized ads for a variety of placements. I was tasked with inpainting Xmas themed ad for a few major singers. The weirdest thing was when the inpainting tool added strange people to an image. This singer was all decked out in tinsel and red, and the inpainting model added a grumpy old man in a top hat. I don't recall clicking the "Add creepy old man" button. At the time this was Stable Diffusion on the backend, run by a variety of model hosting services, Amazon being one. They all had different requirements for the input image and that made things really complex. For some the aspect ratio was impossible to meet, and it would fail if the banner was 200x60. For others, you had to resize it before input, which meant you were adding an image with poor resolution to start. Garbage in, garbage out. All of this to say, there is a lot of preproduction that went into it, and the client never ended up using my attempts.
- giancarlostoro 3mo ago> For others, you had to resize it before input, which meant you were adding an image with poor resolution to start. Thats because small models like SD (Stable Diffusion) are trained on very specific resolutions, its the fancier models that are trained on higher quality, or more diverse sets of resolutions, and if you use a higher quality model to generate lower resolution images, what's actually happening is you're trimming a much bigger image and getting a chunk of it output, at least that's how it feels based on my many hours of experimenting. If I use major models and try to center a thing, I never see it in the center. :) My GPU can only handle so much.
- vunderba 3mo agoSo traditionally, the way you’d do this (and why some UIs like automatic1111 let you configure inpainting so flexibly) is that you didn’t have to shrink the entire image. The general idea was: you mask the area you want changed, and the model inpaints that region at full resolution. The advantage of masking, compared to plain img2img, is that you’re not sending the entire picture to the model. With the classic setups like SD 1.5 and SDXL, you’d effectively inpaint at full resolution: take the masked area from a larger image, scale just that region to the model’s native resolution, process it at the full ~1 megapixel then scale it back and composite it into the original. This lets you add MORE detail. Unfortunately if the OP is using hosted SD models, they might not have that granular control and thus would suffer pretty bad quality loss.
- pattilupone 3mo agoI want a version of this for manga (for translation). Right now I think the go-to lightweight inpainting model for anime and manga is LaMa which is several years old now and it feels like there is room for improvement.
- matthewfcarlson 3mo agoI've been working on trying to outpaint an animated program for my son (Leapfrog Letter Factory if you're curious) and then upscale it. Doing so locally has been actually fairly difficult. I wonder if you could retrain or fine tune this model. They mention building an expert, I wonder if that expert could understand more about translating various characters.
- gifhater 3mo ago[dead]
- simonw 3mo agoI got this working with ONNX (thanks, Claude Opus 4.8) and now I have an interactive demo of the model running entirely in the browser here (~1.3GB download): https://simonw.github.io/moebius-web/ https://simonw.github.io/moebius-web/ - code here: https://github.com/simonw/moebius-web https://github.com/simonw/moebius-web (Claude Code transcript: https://gisthost.github.io/?58039ba5c1ca3ed177e8659168996ee4 https://gisthost.github.io/?58039ba5c1ca3ed177e8659168996ee4) Wrote this up in more detail on my blog: https://simonwillison.net/2026/Jun/22/porting-moebius/ https://simonwillison.net/2026/Jun/22/porting-moebius/
- K0IN 3mo agoAwesome, I wanted to do the exact same thing (used gpt 5.5 + code) but it didn't get the model to work in onnx...
- g58892881 3mo agowell done! unet weights are in fp32. did you by any chance try something lower, fp16?
- da_grift_shift 3mo agoThe model considered it. There are 25 or so mentions of fp16 and fp32 weights across the 7500+ words of Markdown text it generated. So the next question might be: Did it make the right calls? https://github.com/simonw/moebius-web/blob/main/notes.md https://github.com/simonw/moebius-web/blob/main/notes.md https://github.com/simonw/moebius-web/blob/main/plan.md https://github.com/simonw/moebius-web/blob/main/plan.md https://github.com/simonw/moebius-web/blob/main/research.md https://github.com/simonw/moebius-web/blob/main/research.md https://github.com/simonw/moebius-web/blob/main/understanding.md https://github.com/simonw/moebius-web/blob/main/understandin...
- chatmasta 3mo agoWhat is inpainting? Everyone in the comments seems to be familiar with the term, and I don’t see it described in the linked page.
- torgoguys 3mo agoClick on the visualizations to see it in action. The purple areas are areas a user highlighted to tell the system to inpaint, and when you click on the image you see the results of the inpainting. Basically the model redraws sections of an image (the purple areas) using the context of what's in the non-purple areas to decide what might look best in the purple areas. Often used for removing objects but as you can see in the examples it can do other things too.
- NooneAtAll3 3mo ago> and when you click on the image ah, bad UX
- nickandbro 3mo agoHere is a little app I made that allows you to experiment with all of the fine tuned models that runs entirely in your browser: https://inpaintlab.com/ https://inpaintlab.com/
- Zopieux 3mo agoNot great. The inpainted areas are, as usual, very smooth compared to the detailed, "high frequency" look of natural photos. Barely useful enough to erase things in thumbnails.
- vunderba 3mo agoThis and these are cherry-picked examples. The one removing that high tension wire in the nature photo is especially bad. You can literally see the band where it erased it. Even the standard restore tool in Photoshop from years ago can do a comparable job.