5 ms·
Once you have the models on local storage you can move pretty quickly from there to VRAM, I've never found that to be the biggest bottleneck. The problem is pro
by tedivm 1y ago
Once you have the models on local storage you can move pretty quickly from there to VRAM, I've never found that to be the biggest bottleneck. The problem is provisioning itself, especially if you have to actually move models locally. Some of this can be avoided with extremely expensive networking (infiniband to a NAS with model weights), but that's not something you're going to have fun dealing with in a cloud environment.
It might help to remember that the training process is essentially a game of "how fast can we shove data into these GPUs", and having a GPU sit idle because the data can't get into it fast enough is a challenge people have been tackling since at least the P100 series. This has resulted in improvements on the GPUs as well as all the hardware around them. Getting data into the chips is one of the most efficient processes at this point.
- freeqaz 1y agoHow do Serverless GPU Cloud Providers deal with that then? Do they go down the Infiniband-to-NAS rabbit hole to build all of their infrastructure? Or do they just setup an NVME RAID cache to hold the models locally? (Maybe an LRU? With the system memory also being used?) I imagine in the real world that model usage follows a zipfian distribution, ie, a small number of models (<10) represent 95% of the machines for inference workloads. And, for those machines, you can just load the weights off of your ~40gbit ethernet connection since they're never cycling. But for that last 5%, I feel like that's where it becomes important. If I'm running a weird, custom model and I want Lambda-like billing... what's the stack? Is the market big enough that people care? (And do most people just use LORAs which are much easier to hot swap?) Training I imagine is a totally different ballpark because you're constantly checkpointing, transferring data at each step, etc, versus inference. That's a world I know a lot less about though!