8 ms·
1-Bit Bonsai Image 4B Image Generation for Local Devices
- sorenjan 4mo agoThey call it a diffusion model, but it's based on Flux.2 which is a rectified flow model.
- samsartor 4mo agoPersonally I think it's fine to use "diffusion" to refer to the whole family of models
- MitPitt 4mo agoLately I've noticed posts with barely 10 points getting to HN frontpage. Was it always like this?
- Aboutplants 4mo agoI just assume bots
- iamjackg 4mo agoBots doing what? How would the poster being a bot influence why the post itself makes it to the front page with just 10 points?
- speedgoose 4mo agoIt’s about how quickly they get those points. It doesn’t have to be bots. Sending a post to friends with reputable human profiles, and asking for a vote kinda works of most social networks. Some social networks claim they have protection against this but I wouldn’t bet they catch everything.
- DannyPage 4mo agoNot as much competition on the weekend?
- s-macke 4mo agoOn weekends, yes. During the week, that’s also true if they arrive within a short time frame, e.g., three minutes. Almost no one looks at “New”. That is the real issue.
- robbomacrae 4mo agoI believe it's the way the HN algorithm works. In order to give new and obscure posts a shot, it will add them to peoples feeds in their front page and see how they measure. Otherwise new posts wouldn't get seen and the flywheel would never get started. So everyone acts as a sort of beta tester for obscure posts.
- nickvec 4mo agoIf you are looking to see the "true" HN frontpage (i.e. most upvoted posts), I'd recommend using https://hckrnews.com https://hckrnews.com
- moebrowne 4mo agoIf you want a list of posts simply ordered by upvotes: https://news.ycombinator.com/best https://news.ycombinator.com/best
- blurbleblurble 4mo agoMaybe the algorithm has some kind of "momentum" to it, taking into consideration the velocity of upvotes.
- yieldcrv 4mo agoimpressive, combines a couple techniques that I always wanted the frontier models to have having trouble loading the webgl browser demo on my phone but no biggy
- lumost 4mo agoI actually can’t wait for the future where I upgrade hardware in order to upgrade my ai as an alternative to an expensive subscription. There are many problems I want to work on which require billions of tokens. These are completely inaccessible without corporate project sponsorship at the moment. An asic generation machine which can pump out a few 10s of thousands of tokens per second at opus4.6 quality is more than sufficient.
- bigmadshoe 4mo agoCan you give an example of such a problem?
- jjcm 4mo agoDecompiling a binary and recreating the source, doing a full line-by-line security audit, always-on agents monitoring state minute-by-minute, etc. I would very easily find ways to hit that level of token usage if it was cheaper/faster.
- lumost 4mo ago"Design me a 3d printable rocket engine for a hobby rocket project. Verify it's design in a full simulation. Iterate until it works reliably in simulation based on a verified printable design on a consumer laser sintering device (or substitute contract manufacture for under 1000 dollars)." This is a hobby version of a project, but you can imagine commercial versions of the same prompt for new databases, genomics studies, material analysis, operating systems etc.
- SiempreViernes 4mo agoFrom the prompt it seems evident the envisioned user doesn't have an interest in designing the motor themselves, so why not simply buy a stock motor?
- lumost 4mo agoI can't put a 10 page narrative on how my specific motor should work into a hacker news post ;) you can also imagine the above where the goal is to have the ai exceed the performance of stock motors.
- a1o 4mo agoAnyone could pickup the minimal hardware requirements for this? Like both RAM and Storage?
- mkl 4mo agoThe white paper says "mean-active memory pressure down to 1.95 GB for 1-bit Bonsai Image 4B and 2.38 GB for Ternary Bonsai Image 4B". Storage is on the linked page, and is about half that.
- a1o 4mo agoThat is very low, looks like it should run in base MacMini M4 with 16GB RAM. I understand it is not released yet? What sort of harness is necessary for this type of model? (I have only used coding agents through GH Copilot in VS Code, the JetBrains AI tool and Pi, this last one was sort of a pain to setup…)
- smallerize 4mo agoThey are released in the Bonsai Studio software and also https://huggingface.co/collections/prism-ml/bonsai-image https://huggingface.co/collections/prism-ml/bonsai-image
- tcarambat1010 4mo agoFor ternary mlx, size on disk is 3.8GB. 512x512 peak memory use is ~3.7
- SilentM68 4mo agoQuestion, Is it compatible with Ollama, ComfyUI or are those providers unneeded, compatible with low-end hardware? Also, where does "./setup.sh/ drop the components in Linux? Thank you, Sol
- wiradikusuma 4mo agoIs there a benchmark of local image generation models? Local = can run on a 16 GB MacBook or 8 GB+ NVIDIA card.
- liuliu 4mo agoExcept the two (GPT-Image-2 and Nano Banana Pro), anything displayed here can run on the 16 GiB MacBook (including the FLUX.2 [dev]): https://tests.drawthings.ai/generate https://tests.drawthings.ai/generate
- vunderba 4mo agoI run a moderately popular image comparison benchmark site called GenAI Image Showdown [1]. You can click “View All Models” and filter the list down to just locally runnable options (Flux, Qwen, Hunyuan, etc.). https://genai-showdown.specr.net https://genai-showdown.specr.net
- janniks 4mo agoI was expecting to see images of Bonsai trees when I clicked this
- tobr 4mo agoI expected a small tree in black and white pixel art.
- potatoman22 4mo agoI wonder why they didn't use a Bonsai model as the text encoder
- sudb 4mo agoVery interested to see where this kind of work goes for on-device video generation!
- iJohnDoe 4mo agoDoes anyone ever get their stuff to actually work. Like actually load?
- jeroenhd 4mo agoThe online demos require WebGPU so Firefox on mobilr and privacy enhanced browsers will break. WebGPU support on Linux and other open source systems is also trash, you can force it to work in Chrome but it won't be happy.
- petercooper 4mo agoCan't speak for browser demos, but I just got the ternary model working on my M5 generating images. The 1 bit didn't work, as it has a known bug with XCode 24.5 and I wasn't in the mood for installing 24.4 alongside. Here's a generation in your honor: https://peterc.org/img/johndoe.png https://peterc.org/img/johndoe.png
- Havoc 4mo agoYeah worked fine in browser. NVIDIA Card Firefox wayland
- jeroenhd 4mo agoCouldn't try it because the demo app is iOS only and the web version just crashes my browser. The small model is impressive but if you front load a 1.8GB text encoder model, the savings aren't quite as useful. I do wonder how these compare to existing image generation models. I've tried https://github.com/alichherawalla/off-grid-mobile-ai https://github.com/alichherawalla/off-grid-mobile-ai for a while but I find the image generation models rather lacking.
- captainregex 4mo agowhat trade off would one need to clear to justify the hardware and the work to get this running locally as part of a broader system? It’s a lot of work setting up and maintaining a production harness/system on a local device. I don’t personally repeatedly generate images at a scale where using a lab’s app somehow burns all my tokens. I like the ideas of local ai but I don’t see widespread adoption of it happening in commercial or customer situations anytime soon no matter how little/good enough they get. Even Uber- token burn whiplash but I doubt their answer will be “run some of it local”. IT nightmare, I’d imagine.
- smallerize 4mo agoTo our knowledge, Bonsai Image 4B is the first image model in its parameter class to run directly on an iPhone. Isn't SD XL 3.5B? And the refiner model is even larger. Those can run on an iPhone 13 Pro.
- woadwarrior01 4mo agoThe text encoder is still 4-bit quantized.
- mft_ 4mo agoGenuine question: is this solving a real problem? IME, the bottleneck when using diffusion models isn't storage space or memory, it's generation time. Lots of models will run on 8-12 GB 1080-generation GPUs onwards, or on Macs with similar memory, which are probably the bottom end from a GPU power perspective anyway. I also note that these models are marginally slower than the small FLUX.2 model they're based on. Okay, maybe this allows running a local model on something that has a reasonably powerful GPU and limited memory, like an iPhone, but is that really a common requirement?
- soerxpso 4mo agoIt's useful progress. Decent-fidelity local-scale inference means that you can create a product that generates throwaway images frequently without worrying about cost. Thus far every product I've seen that generates images is metered, which severely limits the value. I don't know if this is actually at the "decent fidelity" point yet.
- wmf 4mo agoFor free users, I guess local generation is going to be faster than waiting in a queue.
- moralestapia 4mo agoGenuine question: doesn't it blow your mind that there exists a 1 Gigabyte file/program that can generate any image you can think of just from a rough description of it?
- hk__2 4mo ago> doesn't it blow your mind that there exists a 1 Gigabyte file/program that can generate any image you can think of just from a rough description of it? I can make this into a 5-lines Python program. I’m not saying the images will match the description, but that isn’t part of your spec ;)
- mft_ 4mo agoYeah, it's pretty incredible. And I guess that's mostly what's behind the question: whether this is more of an impressive research/technique demonstrator, or a real product advancement solving a need.
- junto 4mo agoJust a side note, that this website is classified by Apple as an Adult website. I have Limit Adult Websites set in Content & Privacy Restrictions switched on. Led me to wonder what happens if a domain gets a new owner, and they want to petition Apple to remove the block.
- moralestapia 4mo agoThis is why I don't think the big AI companies and nvidia will dominate the market. AIs will just run locally, on whatever hardware you have. Perhaps that's why they worked on this yet-to-be-defined partnership with ARM.
- maephisto666 4mo ago[dead]
- jijji 4mo agoUsing the demo and typing in "A sign that says xxxx" where xxxx is any text, it gets it wrong almost 100% of the time.
- danielEM 4mo agoIs there a way to run it on Vulkan?
- hatch_q 4mo agoNo. Sadly, NVIDIA killed any kind of compute via Vulkan. I took few minutes to try to make it work on ROCm (AMD's alternative to CUDA), landed in python dependency hell.
- WithinReason 4mo ago> NVIDIA killed any kind of compute via Vulkan What do you mean? They are the ones introducing the matmul extensions to Vulkan, which makes compute like this possible
- mk_stjames 4mo agoI saw '1-bit' and my mind first went to 1-bit dithered B&W image generation, not 1-bit model weights.... and so now I'm wondering how cool /fast / compressed a diffusion image generator could be if the images it was trained on / space it worked in was limited to 1 bit (Floyd-Steinberg / Atkinson / your favorite algo here) dithered images. Training would surely be pretty quick and probably fit onto one modern GPU.
- appplication 4mo agoThis was exactly where my mind went as well and I think there would be some really cool ideas to explore here
- Retr0id 4mo agoI think you'd still be better off training in greyscale and dithering after the fact.
- ttul 4mo agoWithin a day, someone will have trained a LoRA for this 1-bit model that enables hentai content generation on your Apple Watch.
- Zopieux 4mo agoGreat.
- liuliu 4mo ago> To our knowledge, Bonsai Image 4B is the first image model in its parameter class to run directly on an iPhone. This is wrong. But they worded it carefully to be not entirely wrong. FLUX.2 [klein] 4B (the same parameter class, basically the same model) runs on iPhone through Draw Things app, with 8-bit or 6-bit quantization (hence not "directly", I guess, but that is the technicality that sounds fishy enough).
- dbcooper 4mo agoA few implementations listed on LM Studio. Any recommendations for which one to use?
- cadamsdotcom 4mo agoStuff like this is great - more promises of things that can run on phones please! Sadly right now the expensive developer subscription means the few folks willing to hold a forever subscription make something that barely works then move on… or make something with so many ads it is an app. For example Google’s “Model Garden” app has no ads but still has major UX issues and isn’t suitable for daily use, even though the models are amazing. Raising awareness of how capable today’s phone hardware is will make normal people demand to run what they choose on their phones. It’d be a much stronger way back to general purpose computing than via all legislation that has been tried so far..
- huflungdung 4mo ago[dead]
- flashman 4mo agoTwenty years ago, I don't think any of us were excited about a future internet where we couldn't trust whether what we were seeing or reading was genuine. I hope one day we'll be able to look back on this era as an aberration, like that scene in Mad Men where the Drapers fling their picnic rubbish onto the grass and drive away.
- Cider9986 4mo agoWhat could make it stop?
- nostrebored 4mo agoI am pretty excited. The factuality of important events has been distorted for most of history. Moving to a low information trust society is something that I think will be positive.
- mungoman2 4mo agoCurious about this take, how do you mean? I understand the point of distorted facts, but what I’m not sure how things are improved by basically having no trust in any facts?
- enduser 4mo agoI’m not the original poster, but I think they are saying if we are more skeptical about what we read, we are less likely to absorb propaganda as fact.
- sroussey 4mo agoI extracted the code from the web demo to add to make a web image generation node to my in browser ai workflow tool, and it’s pretty sweet. Waiting for xenova to add to transformersjs 4.3 and I’ll release as well. Couldn’t wait though to test.
- lwansbrough 4mo agoCan anyone think of any negative externalities of making generative photorealistic images illegal? I can think of a lot of positives. The negatives amount to a convoluted argument about the limits of free speech.
- IncreasePosts 4mo agoPrisoner 1: so, what are you in for? Prisoner 2: I made a picture of a nice sunset over the ocean
- lwansbrough 4mo agoIf it were illegal it wouldn’t be readily available. You’d have to seek it out. People seeking it out wouldn’t be using it to generate a sunset.
- ToValueFunfetti 4mo agoIf it were illegal to generate a sunset, people would absolutely seek it out to generate a sunset, if just as an act of civil disobedience. Look at how people reacted to being told sharing a number is illegal[1]. Why would they act differently about an illegal computation? [1]https://en.wikipedia.org/wiki/AACS_encryption_key_controversy https://en.wikipedia.org/wiki/AACS_encryption_key_controvers...
- kordlessagain 4mo agoI've tested this and it's not as good as Flux in my opinion.
- kordlessagain 4mo agohttps://github.com/kordless/bonsai-docker https://github.com/kordless/bonsai-docker if you want to run without fiddling with the local filesystem.
- deleted 4mo ago[deleted]
- Songjinhao 4mo ago[flagged]
- edf13 4mo agoOdd… UK visitor and I get: Website Not Allowed “prismml.com” is a restricted website.
- n3xyf 4mo agoThis is cool and all but is there a real use case for these? One that actually creates value?
- hmokiguess 4mo agoGot it to run on iPhone but was surprised to see they have some form of censorship and moderation on the input side on their client app. I thought a big part of local/offline AI was sovereignty, unfiltered, and censorship/bias resistance.
- jermaustin1 4mo agoIt looks like the text encoder is a separate bonsai model, so it can probably be obliterated or whatever it is called.
- aalam 4mo agoAbliteration. The word and its descriptions read like pure sci-fi https://huggingface.co/blog/mlabonne/abliteration https://huggingface.co/blog/mlabonne/abliteration
- vorticalbox 4mo agothey have a webGPU demo [0] at 4 steps it takes 7 seconds to generate an image on my M4 https://huggingface.co/spaces/webml-community/bonsai-image-webgpu https://huggingface.co/spaces/webml-community/bonsai-image-w...
- baisampayans 4mo ago[flagged]
- willXare 4mo ago[flagged]
- deleted 4mo ago[deleted]
- willXare 4mo ago[flagged]
- willXare 4mo ago[flagged]