3 ms·
Quantization in this context is the precision of each value in the vector or matrix/tensor. If the model in question has a token embedding length of 1024, even
by Acumen321 3y ago
Quantization in this context is the precision of each value in the vector or matrix/tensor.
If the model in question has a token embedding length of 1024, even if it was a 1 bit quantization, each token has 2^1024 possible values.
If the context length is 32,000 tokens, there are 32,000^2^1024 possible inputs.