4 ms·
> Google trained Llama Source? This would make quite the splash in the market
by DrBenCarson 2y ago
> Google trained Llama
Source? This would make quite the splash in the market
- xnx 2y agoIt's in the article: "When training the Llama-2-70B model, our tests demonstrate that Trillium achieves near-linear scaling from a 4-slice Trillium-256 chip pod to a 36-slice Trillium-256 chip pod at a 99% scaling efficiency."
- llm_nerd 2y agoI'm pretty sure they're doing fine-tune training, using Llama because it is a widely known and available sample. They used SDXL elsewhere for the same reason. Llama 2 was released well over a year ago and was training between Meta and Microsoft.
- hhh 2y agoThey can just train another one.
- llm_nerd 2y agoLlama 2 end weights are public. The data used to train it, or even the process used to train it, are not. Google can't just train another Llama 2 from scratch. They could train something similar, but it'd be super weird if they called it Llama 2. They could call it something like "Gemini", or if it's open weights, "Gemma".
- lern_too_spel 2y agoThe article says they used maxtext to load the weights and pretrain on additional data. It looks like the instructions for doing that are here: https://github.com/AI-Hypercomputer/maxtext/blob/main/getting_started%2FRun_Llama2.md https://github.com/AI-Hypercomputer/maxtext/blob/main/gettin...
- ein0p 2y agoThey don't mean literally LLaMA. They mean a model with the same architecture.