4 ms·Plus all the flexibility of CUDA is not needed for LLMs anywaysby WanderPanda 2y agoPlus all the flexibility of CUDA is not needed for LLMs anyways