4 ms·
For production loads? No one will run DS without expert parallelism so it’s not important for model to fit on one gpu.
by boroboro4 2mo ago
For production loads? No one will run DS without expert parallelism so it’s not important for model to fit on one gpu.
- smallerize 2mo agoDo you mean people will run multiple GPUs, or that streaming a small % of experts from disk won't completely kill performance?
- touisteur 2mo agoTime for AMD to improve their directstorage story I guess. Is there a good writeup on how one would use it there ? Same for gpudirect (more useful for scale-out or training). Everyone is focused on AI but these are two interesting techs trickling down from the NVIDIA tree, relatively "easy" to use there, which maybe exist in AMD world but I somehow missed the docs and APIs and demos on how to use them...