4 ms·
That’s the size of the model. Running it requires additional space for the KV cache depending on how much context you use. You can probably get it running in 8G
by Aurornis 8d ago
That’s the size of the model. Running it requires additional space for the KV cache depending on how much context you use. You can probably get it running in 8GB of RAM for short outputs but to get okay context length you’d want more.