3 ms·
Let’s see, so if you get the same 1/9th the size compression ratio with GLM-5.3-Flash, then you’d end up with a ~72GB model that’s about as good as GPT-5.6 Sol
by Chance-Device 16d ago
Let’s see, so if you get the same 1/9th the size compression ratio with GLM-5.3-Flash, then you’d end up with a ~72GB model that’s about as good as GPT-5.6 Sol (high), according to artificialanalysis.ai
Which is within reach of some higher end consumer hardware, especially with layer offloading.
You have to wonder what kind of trouble the “labs” are in when this is becoming possible. Lots of money, where’s the moat?
- indy 16d agoThey're trying to build a moat with legislation, using fear over 'safety' as an excuse to ban these open models
- DoctorOetker 15d agoA gentleman's agreement between the western bloc, China, India, Russia, ... ? It would be a lip service agreement, and all would continue the machine learning race... The mere suggestion of slowing down AI progress signals to adversaries that there is some novel power just identified: do you think adversaries would agree and slow down, or calculate a little harder and longer in order to also figure out what it unlocks?
- indy 14d agoDo you ever wonder how Anthropic or OpenAI can release a new state of the art model and then a few months later there's a Chinese open model that's nearly as good?
- kllrnohj 16d agoThe labs still have performance as a differentiator for coding usages and similar, and for other things there's still all the same reasons people switched cloud hosted stuff in the first place. AWS & friends didn't get popular because the hardware was out of reach, after all.
- Chance-Device 16d agoThat’s an argument for a cloud LLM service, sure, but the hyperscalers can do that by themselves with the weights. What’s the moat for trillion dollar AI companies? Access-anywhere convenience for models as good as everyone else’s?
- kllrnohj 16d agothe trillion dollar AI companies will be the chip designers for the hyperscalers, and probably not worth a trillion dollars as a result
- ctolsen 16d agoIf we go with AA's benchmarks Qwen 3.8 27B is already slightly below Luna level which is in itself impressive, but with this compression it should be just slightly more below Luna level and could run on my old GTX 1070 that I'm now tempted to fire up. That's kinda nuts even allowing for small-model problems that I'm sure I'd see quite clearly.