2 ms·
Mind sharing your setup? I also have dual 3090s, but getting nowhere close to 300k context limits with 4 bit quantized models at that size (using vllm).
by l332mn 4mo ago
Mind sharing your setup? I also have dual 3090s, but getting nowhere close to 300k context limits with 4 bit quantized models at that size (using vllm).