3 ms·
https://www.techpowerup.com/gpu-specs/geforce-rtx-4090.c3889 https://www.techpowerup.com/gpu-specs/geforce-rtx-4090.c3889 Memory Size: 24 GB Memory Typ
by schmidtleonard 2y ago
https://www.techpowerup.com/gpu-specs/geforce-rtx-4090.c3889 https://www.techpowerup.com/gpu-specs/geforce-rtx-4090.c3889
Memory Size: 24 GB
Memory Type: GDDR6X
Memory Bus: 384 bit
Bandwidth: 1.01 TB/s
Bandwidth between where the LLM is stored and where your matrix*vector multiplies are done is the important figure for inference. You want to measure this in terabytes per second, not gigabytes per second.
A 7900XTX also has 1TB/s on paper, but you'll need awkward workarounds every time you want to do something (see: article) and half of your workloads will stop dead with driver crashes and you need to decide if that's worth $500 to you.
Stacking 3090s is the move if you want to pinch pennies. They have 24GB of memory and 936GB/s of bandwidth each, so almost as good as the 4090, but they're as cheap as the 7900XTX with none of the problems. They aren't as good for gaming or training workloads, but for local inference 3090 is king.
It's not a coincidence that the article lists the same 3 cards. These are the 3 cards you should decide between for local LLM, and these are the 3 cards a true competitor should aim to exceed.
- Dylan16807 2y agoA 4090 is not "years old pleb tier". Same for 3090 and 7900XTX. There's a serious gap between CXL and RAM, but it's not nearly as big as it used to be.