3 ms·
https://github.com/NVIDIA/TransformerEngine https://github.com/NVIDIA/TransformerEngine - mixed precision FP8 is already here and provides similar accuracy to F
by eth-mld 3y ago
https://github.com/NVIDIA/TransformerEngine https://github.com/NVIDIA/TransformerEngine - mixed precision FP8 is already here and provides similar accuracy to FP16/BF16
- villgax 3y agoYeah and for some reason limited to only hardware support of H100. Even the cost of doing it in software is outweighed by the speed & storage gains from it