4 ms·
I am amazed, though not entirely surprised, that these models keep getting smaller while the quality and effectiveness increases. z image turbo is wild, I'm loo
by codezero 9mo ago
I am amazed, though not entirely surprised, that these models keep getting smaller while the quality and effectiveness increases. z image turbo is wild, I'm looking forward to trying this one out.
An older thread on this has a lot of comments: https://news.ycombinator.com/item?id=46046916 https://news.ycombinator.com/item?id=46046916
- roenxi 9mo agoThere are probably some more subtle tipping points that small models hit too. One of the challenges of a 100GB model is that there is non-trivial difficulty in downloading and running the thing that a 4GB model doesn't face. At 4GB I think it might be reasonable to assume that most devs can just try it and see what it does.
- Auracle 9mo agoQuality is increasing, but these small models have very little knowledge compared to their big brothers (Qwen Image/Full size Flux 2). As in characters, artists, specific items, etc.
- efskap 9mo agoI smell the bias-variance tradeoff. By underfitting more, they get closer to the degenerate case of a model that only knows one perfect photo.
- vunderba 9mo agoAgreed - given what Tongyi-MAI Lab was able to accomplish with a 6b model - I would love to see what they could do with something larger. Somewhere in the range of 15-20b, between these smaller models (ZiT, Klein) and the significantly larger models (Flux.2 dev).
- littlestymaar 9mo agoThat's what LoRAs are for. And small models are also much easier to fine tune than large ones.
- Auracle 9mo agoI hate that excuse. I want the model to know who the Paw Patrol is without either finding a lora (which probably won't exist because they're mostly porn) or needing to make a dataset, tag it, and then train it myself.
- aitchnyu 9mo agoIs there a theoritical minimum for params for a given output? I saw news about GPT 3.5, then Deepseek training models at a fraction of that cost, then laptops running a model that beats 3.5. When does it stop?