3 ms·
I've actually done some work and was able to get the full 12GB Z-image Turbo to run on my 8Gb 3070, and throughout the generation only 600MB of VRAM was allocat
by YuechenLi 18d ago
I've actually done some work and was able to get the full 12GB Z-image Turbo to run on my 8Gb 3070, and throughout the generation only 600MB of VRAM was allocated (you can get a speed-up by double buffering, but the maximum VRAM was still only ~1GB through inference). It's very experimental and pretty much requires you to write all the compute kernels directly specifically against that particular weight and statically allocate the memory at compile time instead of using current ML frameworks like Pytorch/JAX, but I think in time somebody else would figure it out.
- ASalazarMX 17d agoThe hardware savings at the scale Anthropic or Google use would me immense, it makes me wonder why no big player has done more optimization already. When DeepSeek showed how inneficient were the models of its time, I'd expected each to create a permanent optimization team with all the talent they have hired. I guess hardware is not that expensive to them in the grand scheme of things, at least not at this stage. OTOH, their propietary models might be thoughly optimized and we can't know, because they're still bound by supply contracts to buy the same amount of hardware nevertheless.