5 ms·
2x RTX 3090 is not enough for proper use. ideal is 64GB+ so that you get proper cache and 256k context with good enough models. (e.g. Qwen 3.8 27B with MXFP4).
by nicce 27d ago
2x RTX 3090 is not enough for proper use. ideal is 64GB+ so that you get proper cache and 256k context with good enough models. (e.g. Qwen 3.8 27B with MXFP4).
And that is tight already.
You would need 3x RTX 3090 - but then the PCIe bandwidth comes an issue if you really want tensor parallelism with three cards. 3x PCIe5 x16 is not cheap with direct CPU access.
So surprisingly, 2x r9700 starts be a nice deal.
> Also these 20-30t/s jumping to 150-200... Watch out for the massaged numbers coming from vendors.
Well, luckily these are not vendor numbers. Prefill also scales almost linearly with the amount of GPUs.
- formerly_proven 27d agoYou don't need PCIe 5.0 x16 since RTX 30 are not PCIe 5.0 to begin with.
- nicce 27d agoWell, that makes them just even slower then
- snovv_crash 27d agoThe whole point of dual 3090 over other cards is the nvlink support. At that point the pcie doesn't really matter.
- nicce 27d agoOops
- latentsea 26d ago> 2x RTX 3090 is not enough for proper use I don't see how. Even on 32GB you can run Q6_K_XL quant with MTP at 200k context k=q8_0, v=q5_1. So 48GB VRAM is good enough to run Q8 at long context. Also with tings like ninfer and it's various forks I'm seeing people get very good performance out of Qwen3.8 models on all sorts of NVIDIA cards.