2 ms·
Curious which models are you able to run and how many 3090s do they require at scale?
by vladgur 5mo ago
Curious which models are you able to run and how many 3090s do they require at scale?
- mips_avatar 5mo ago4 3090s with nvlinks on each pair. Super fast inference on Moe models around 20-36b
- embedding-shape 5mo ago> Super fast inference How fast is "super fast" exactly, and with what runtime+model+quant specifically? Curious to see how how 4x 3090s compare to 1x Pro 6000, could probably put together 4x 3090s for a fraction of the cost compared to the Pro 6000, but the times I've seen the tok/s in/out for multiple GPUs my heart always drops a little.
- mips_avatar 5mo agoI haven't benchmarked against a pro 6000, it's more that i have 4 3090s and i don't have a pro 6000.
- embedding-shape 5mo agoYes, that's why I'm asking you what exactly 4 3090s get in prompt-processing and generation, sorry if I was unclear.
- mips_avatar 5mo agoMaxes out around 4K tok/s output. Each pair of 3090s has its own instance of the model with parallelism across the nvlink bridge. Though nvlink is only 2x over pcie5