3 ms·
Where does this 3% figure come from?
by joaogui1 3y ago
Where does this 3% figure come from?
- Zuiii 3y agoThese models are generally trained on tarabytes of data, but are usually 10s of gigabytes large (or much less if quantized). The latest true open source model, mistral 7b, is only 3GB (0.3% of a TB) when quantized.
- rpdillon 3y agoI did a very similar analysis with Llama 65B being trained on 5.6T tokens assuming token length of 4 characters and comparing with a quantized model size of ~38GB. The 3% number was a conservative rounding of the same calculation, but retaining fp16 rather than quantizing to 4 bits. Here's my original back of the napkin analysis: https://news.ycombinator.com/item?id=36681440 https://news.ycombinator.com/item?id=36681440