2 ms·
L40S sound good on paper, but the memory bandwith compared to even a 40G A100 is reduced in half which is crazy in the era when we are bottlenecked by it rather
by treesciencebot 3y ago
L40S sound good on paper, but the memory bandwith compared to even a 40G A100 is reduced in half which is crazy in the era when we are bottlenecked by it rather than the actual compute. It costs as much (or even more) than an 80G A100, but instead of getting ~2TB/s, you get ~800GB/s.
- treesciencebot 3y agoL40S might be a more appealing choice if you are switching from an A40, but then why the hell aren't you using A100 which is actually much cheaper and much more commonly available. The only perceivable reason I can see is fp8 support, but even with it, I don't think it is worth the price.
- YetAnotherNick 3y agoYour statement is only partially correct. For single batch inference, something we do locally, the bottleneck is generally bandwidth specially for llm. But no one does single batch inference in server. Also training is bottlenecked by compute and not memory bandwidth.