3 ms·
OpenLLaMA 7B Training Completed to 1T Tokens
- fancyfredbot 3y agoThis is great. Based on the throughout of 2200 tokens/sec and the 1,000,000,000,000 tokens used to train this was at least $183k worth of compute (that's based on the three year committed use rate). And now we can have it for free!
- thawab 3y agoThe price for training their 7B, as stated by MosaicML[0] and Falcon 7B, is roughly the same. [0] https://twitter.com/MosaicML/status/1660738892306485248 https://twitter.com/MosaicML/status/1660738892306485248
- mdaniel 3y agobe sure to read the warning in their repo: https://github.com/openlm-research/open_llama#loading-the-weights-with-hugging-face-transformers https://github.com/openlm-research/open_llama#loading-the-we... > Please note that it is advised to avoid using the Hugging Face fast tokenizer for now, as we’ve observed that the auto-converted fast tokenizer sometimes gives incorrect tokenization