3 ms·
Does JAX have its own implementations of matmul, flash attention etc? Or does it use the ROCm implementations like PyTorch does? (e.g,. hipblaslt, Composable Ke
by anthonix1 2y ago
Does JAX have its own implementations of matmul, flash attention etc? Or does it use the ROCm implementations like PyTorch does? (e.g,. hipblaslt, Composable Kernel FA etc)
Not too familiar with JAX, but the abysmal PyTorch training perf on MI300x is in large part attributable to the slow perf of the ROCm libraries it is using under the hood.
- jdeaton 2y agoJAX has a sub-system called Pallas[1] with a Triton-like programming model and an example implementation of Flash Attention [2]. It is quite fast. On TPUs I've heard that the XLA compiler already emits a flash-attention-like computation graph for a regular JAX implementation of attention so there's no need to have some specialized kernel in that case. 1. https://jax.readthedocs.io/en/latest/pallas/index.html https://jax.readthedocs.io/en/latest/pallas/index.html 2. https://github.com/jax-ml/jax/blob/main/jax/experimental/pallas/ops/gpu/attention.py https://github.com/jax-ml/jax/blob/main/jax/experimental/pal...