3 ms·
Do we know if the full model is FP8 or FP16/BF16? The hugging face page says BF16: https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct https://huggingfac
by vFunct 1y ago
Do we know if the full model is FP8 or FP16/BF16? The hugging face page says BF16: https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct
So likely it needs 2x the memory.
- danielhanchen 1y agoI think it's BF16 trained then quantized to FP8, but unsure fully - I was also trying to find out if they used FP8 for training natively!
- jychang 1y agoQwen uses 16bit, Kimi and Deepseek uses FP8.
- danielhanchen 1y agoOh ok cool thanks!