4 ms·
This is about the price of 3 5090's, but it can run the full 670B model with Q8 quantitization, albeit at a few tokens per second. Doing this on GPUs would requ
by lodovic 2y ago
This is about the price of 3 5090's, but it can run the full 670B model with Q8 quantitization, albeit at a few tokens per second. Doing this on GPUs would require 24 5090's (if nvlink still works, but I don't think so) or 10 H100's.