4 ms·
That's 90% bandwidth efficiency and 60% compute efficiency https://www.nvidia.com/en-us/data-center/h100/ https://www.nvidia.com/en-us/data-center/h100/
by WithinReason 2y ago
That's 90% bandwidth efficiency and 60% compute efficiency
https://www.nvidia.com/en-us/data-center/h100/ https://www.nvidia.com/en-us/data-center/h100/
- helloericsf 2y agoThey don't have h100. wink,wink.
- rfoo 2y agoThey have H800s which have exactly same memory bandwidth and max FLOPS.
- pk-protect-ai 2y agoWhat about NVLink? Does it plays a role here?
- rfoo 2y agoFor FlashMLA? No. The code here runs on one GPU only and do not have a builtin communication part.
- pk-protect-ai 2y agoBut for the training it does. You need to communicate gradient changes between GPUs.
- deleted 2y ago[deleted]