3 ms·
Python is not really the bottleneck in LLM applications. It is for tabular RL, but certainly not for deep RL (i have had discussions with DM folk over this in r
by PartiallyTyped 3y ago
Python is not really the bottleneck in LLM applications. It is for tabular RL, but certainly not for deep RL (i have had discussions with DM folk over this in r/RL, and the ppl from stable diffusion).
The problem is the bus, cuda, and the sheer volume of data that need to be transferred.
Pytorch itself is actually a wrapper around torchlib, which is written in C++.
The compilation step of PyTorch 2.0 provides a sizeable improvement, but not 2 orders of magnitude as you’d expect from python to c++ migrations. The compilation is due to the backend more so than python itself. See Triton for example.