3 ms·
Could you educate me on what's really different about a card that is "server-grade" vs. one that is not?
by TallGuyShort 4y ago
Could you educate me on what's really different about a card that is "server-grade" vs. one that is not?
- jmalicki 4y agoTwo of the biggest differences are ECC memory and virtualization support.
- crabbone 4y agoYeah, also while ECC is mostly transparent to the software layer, virtualization support is in itself a huge software layer that comes with this kind of equipment. MIG ( https://docs.nvidia.com/datacenter/tesla/mig-user-guide/ https://docs.nvidia.com/datacenter/tesla/mig-user-guide/ ), there's also software made to expose GPU inside containers. This software is mostly supplied by vendor, and it mostly cares about vendor's tech, and will usually require special features exposed by the hardware, so, won't work with consumer-grade GPUs.
- m463 4y agowell, and it's a specialized die. A100 ~ 2x size of 3090
- ericpauley 4y agoFor these GPUs, it partially comes down to form factor (being able to drop 8 in a chassis with unified cooling). I’d say that’s the primary difference here. This applies maybe a little more noticeably to the rest of the computer around the a100s, with ECC ram, hot swappable components, and redundant power supplies. The 8x a100 rigs are absolute beasts, difficult to move even by two people without removing some components. There is also price discrimination here where manufacturers aim for higher margins on server components (as nvidia clearly does). Between the two mentioned cards, though, it really does come down to performance. The cards are simply not interchangeable for large scale model training.