3 ms·
Depends on your pcie connection. If they're both x16 then it's pretty low overhead, x8 is ok, but x4 is too slow. Also it's a bit tricky getting an optimal se
by mips_avatar 2mo ago
Depends on your pcie connection. If they're both x16 then it's pretty low overhead, x8 is ok, but x4 is too slow. Also it's a bit tricky getting an optimal setups with mismatched vram, I think you could probably still make use of the full vram if you're clever but it's trickier.
- nullc 2mo agofor layer parallelism (e.g. to get more vram) the bandwidth between layers is essentially nothing (like 16kb per token I think), so I don't think x4 would even be a problem!
- ericd 2mo agoGood point. It's much more of an issue when running dense models with tensor parallelism. In that case, I'd look for an MoE model instead.
- deleted 2mo ago[deleted]