3 ms·
I wonder if you can't do that LSH trick to turn it into a sparse matrix problem and run it on CPU that way.
by guywhocodes 4y ago
I wonder if you can't do that LSH trick to turn it into a sparse matrix problem and run it on CPU that way.
- nmfisher 4y agoThat's pretty much what SLIDE [0] does. The driver was achieving performance parity with GPUs for CPU training, but presumably the same could apply to running inference on models too large to load into consumer GPU memory. https://github.com/RUSH-LAB/SLIDE https://github.com/RUSH-LAB/SLIDE