3 ms·
Even that is being tackled by newer GPU architectures. For example, novelai is currently training an LLM in fp8 precision, using H100 GPUs.[1] [1] https://blog
by Jackson__ 3y ago
Even that is being tackled by newer GPU architectures. For example, novelai is currently training an LLM in fp8 precision, using H100 GPUs.[1]
[1]
https://blog.novelai.net/anlatan-acquires-hgx-h100-cluster-4b7a2e6a631e https://blog.novelai.net/anlatan-acquires-hgx-h100-cluster-4...
https://blog.novelai.net/text-model-progress-is-going-good-82a94855445e https://blog.novelai.net/text-model-progress-is-going-good-8...
- superkuh 3y agoCool stuff. I looked at https://en.wikipedia.org/wiki/Hopper_%28microarchitecture%29 https://en.wikipedia.org/wiki/Hopper_%28microarchitecture%29 and I noticed that that the fp8 support is only for the tensor cores and not the CUDA side. Does that mean training with H100 GPU in fp8 mode would use some software ecosystem that's not the existing vast existing CUDA one? Or am I just misunderstanding CUDA cores vs tensor cores? PS, as a joke, they should implement GPU fluint8 and get baked in non-linearity for the activation function without even using a non-linear function, https://www.youtube.com/watch?v=Ae9EKCyI1xU https://www.youtube.com/watch?v=Ae9EKCyI1xU ("GradIEEEnt half decent: The hidden power of imprecise lines" by suckerpinch)
- darknoon 3y agoyou can access the tensor cores from cuda, in practice you might generate the code with something like openai triton