3 ms·
This calculator estimates the GPU memory needed to run LLM inference. Select the model size and precision (FP32 - FP4) to get a quick memory range estimate.
by javaeeeee 1y ago
This calculator estimates the GPU memory needed to run LLM inference. Select the model size and precision (FP32 - FP4) to get a quick memory range estimate.