3 ms·
GLM 5.2 awq4 via Opencode, it was better than corporate's fave Sonnet 4.8 and a but worse than Opus 4.8. overall quite capable of a lot of what the dev team nee
by Y_Y 2mo ago
GLM 5.2 awq4 via Opencode, it was better than corporate's fave Sonnet 4.8 and a but worse than Opus 4.8. overall quite capable of a lot of what the dev team needed. Not as good as Sonnet 5 but also never runs out of tokens.
The flip side is that it takes four H200s to run, and that will only let you cache context for maybe three users.
Fingers crossed our Blackwells show up and Kimi 3 really releases weights, because at some point devs are spending a significant portion of their salary on tokens and it's somehow cheaper to buy these ridiculous DGX servers and rack and run them.
- nojs 2mo agoWhat hardware are you expecting to run K3 on?
- Y_Y 2mo agoEight B200s, at whatever quantization I can fit, ideally fp4. Apparently it might fit in eight H200s, but with 2.8B parameters (1.8% active) I'm not optimistic.
- foobar10000 2mo agoGlm 5.2 nvfp4 on 4 b300 with dpattn 4 and ram will get you about 20 users live at 400k context - and 60 easily if you give that server 2tb of ram and 4 nvme 8 tb drives. There are some sglang patches needed - but we will be releasing them soon.