4 ms·
How about reloading parts of the model as the inference progresses instead of splitting it into GPU/CPU parts? Reloading would be memory-limited to the largest
by bitL 3y ago
How about reloading parts of the model as the inference progresses instead of splitting it into GPU/CPU parts? Reloading would be memory-limited to the largest intermediate tensor cut.
- moffkalast 3y agoThe Tensor Reloaded, starring Keanu Reeves
- regularfry 3y agoThat would turn what's currently an L3 cache miss or a GPU data copy into a disk I/O stall. Not that it might not be possible to pipeline things to make that less of a problem, but it doesn't immediately strike me as a fantastic trade-off.
- bitL 3y agoOne can keep all tensors in the RAM, just push whatever needed to GPU VRAM, basically limited by PCIe speed. Or some intelligent strategy with read-ahead from SSD if one's RAM is limited. There are even GPUs with their own SSDs.