3 ms·
> A system based on 4x RTX6K can run GLM 5.3 at NVFP4 precision It actually runs fine at FP8 on this hardware too, with the full 1M context.
by nojs 24d ago
> A system based on 4x RTX6K can run GLM 5.3 at NVFP4 precision
It actually runs fine at FP8 on this hardware too, with the full 1M context.
- CamperBob2 23d agoFlash will run on 4x, but at the time I ran that test there were no 4-card quants for the full 744B-A40B 5.3 model. There are now, though, with KLD figures close to the FP8 level. I need to do some more benchmarking to see if they live up to the hype.
- nojs 23d agoOh right, I was referring to flash. I haven’t tried these either, but the ones for 5.2 looked interesting.