5 ms·
WebGPU has no equivalent to tensor cores to my understanding; are there plans to add something like this? Or would this be "implementation sees matmul-like code
by why_only_15 3y ago
WebGPU has no equivalent to tensor cores to my understanding; are there plans to add something like this? Or would this be "implementation sees matmul-like code; replaces with tensor core instruction". For optimal performance, my understanding is that you need tight control of e.g. shared memory as well -- is that possible with WebGPU?
On NVIDIA GPUs, flops without tensor cores are ~1/10th flops with tensor cores, so this is a pretty big deal for inference and definitely for training.
- raphlinus 3y agoShared memory, yes, with the goodies: atomics and barriers. We rely on that heavily in Vello, so we've pushed very hard on it. For example, WebGPU introduces the "workgroupUniformLoad" built-in, which lets you broadcast a value to all threads in the workgroup while not introducing potential unsafety. Tensor cores: I can't say there are plans to add it, but it's certainly something I would like to see. You need subgroups in place first, and there's been quite a bit of discussion[1] on that as a likely extension post-1.0. [1]: https://github.com/gpuweb/gpuweb/issues/3950 https://github.com/gpuweb/gpuweb/issues/3950
- jdashg 3y agoYes, we expect to have a natural path towards explicit cooperative matrix multiply ops. If you have a wishlist, we have an issue tracker! ;) https://github.com/gpuweb/gpuweb/issues https://github.com/gpuweb/gpuweb/issues