3 ms·
What hardware would you even be able to run this on?
by unleaded 2mo ago
What hardware would you even be able to run this on?
- adrian_b 2mo agoEven the lowliest hardware could run this, but at an unlikely to be useful low speed, e.g. of 3 or 4 tokens per minute (by reading the weights from a couple of 4 TB SSDs for the BF16 model, or from a 4 TB SSD for the FP8 variant). The question about LLMs is never whether they can be run, because that has a trivial answer, they can always be run. The right question is what speeds are achievable for representative hardware configurations. At launch, it is difficult to estimate the speed. That should be known after someone reports experimental results. Moreover, for many LLMs the speed improved sometimes later after their release, after tweaks in inference backends, like llama.cpp or vLLM.
- badcafe23423435 2mo agotell me how run 35B on my 8GiB VRAM (linux) speed is not problem when You run agents and forget for 2-3 days