3 ms·
If the model is stored in 32bits, it will be ~4x param size You can see the actual size from hugging face. For example, https://huggingface.co/WizardLM/WizardL
by notpublic 3y ago
If the model is stored in 32bits, it will be ~4x param size
You can see the actual size from hugging face. For example, https://huggingface.co/WizardLM/WizardLM-7B-V1.0/tree/main https://huggingface.co/WizardLM/WizardLM-7B-V1.0/tree/main
Size with quantization, https://github.com/ggerganov/llama.cpp#quantization https://github.com/ggerganov/llama.cpp#quantization