7 ms·
> The training for Phi-2 took 14 days on 96 A100 GPUs This would mean that it costs around ~30k USD to train. If training an LLM becomes cheaper than buying
by duchenne 3y ago
> The training for Phi-2 took 14 days on 96 A100 GPUs
This would mean that it costs around ~30k USD to train.
If training an LLM becomes cheaper than buying a car, it could democratize AI a lot.
- eternauta3k 3y agoYou don't need to train it again, Microsoft already did. Unless you want to develop a new one, then you also need the team of researchers/engineers.
- alecco 3y agoNote the model is trained on data generated by GPT-4. It's probably orders of magnitude more expensive to generate the data at current API prices. The whole point of these papers is that training data quality is key. I would much prefer for these companies to release the training data than the weights. But that will never happen. "We speculate that the creation of synthetic datasets will become, in the near future, an important technical skill and a central topic of research in AI."
- verdverm 3y agoThis sounds like the methodology from "Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes" i.e. master teaches apprentice or LLM trains SLM https://arxiv.org/abs/2305.02301 https://arxiv.org/abs/2305.02301 (May '23)
- eightysixfour 3y agoYes, I think we are seeing the beginning of a feedback loop where we can use current LLMs to generate better datasets at a scale large enough to create new LLMs. This is the positive feedback loop that I think is going to make the biggest difference in model quality over the next few years.
- digdugdirk 3y agoWhat do you see as the limit to this improvement?
- eightysixfour 3y agoThere is probably some limit where making the dataset larger, with more diverse information, does not create meaningful improvements with current architectures. I do not know what that limit is or what it looks like, but I also don’t think we are particularly close to it yet. “The Pile” dataset is the asset we needed to jumpstart this process, it had so much raw data it could get us over the hump, but Phi and some of the models trained on explicit reasoning make the limitations of random shit people say on the internet pretty clear.
- pmb22 3y agoThe eightysixfour rule? You would think that this would follow something similar to Moore's law for a little while
- verdverm 3y agoThe Pile dataset for those interested https://pile.eleuther.ai/ https://pile.eleuther.ai/ https://arxiv.org/abs/2101.00027 https://arxiv.org/abs/2101.00027 I'm bullish on domain specific models that start from generalized models. Something of a T shape analogy, but maybe a couple of distillation & fine-tuning steps
- lukeplato 3y agomodels trained on gpt output might be more distilled and specialized but it wouldn't be improving generalization
- lukeplato 3y agohttps://twitter.com/pfau/status/1674766269113937920 https://twitter.com/pfau/status/1674766269113937920
- 3y ago
- IanCal 3y ago> Note the model is trained on data generated by GPT-4. Is it? I couldn't find that in the page, and can't easily access the links. The previous paper used 1B tokens from GPT-3.5 > It's probably orders of magnitude more expensive to generate the data at current API prices. If you're generating a billion tokens, you might do better with dedicated instances, iirc they used to say if you were doing more than a few hundred million a month dedicated things were cheaper.
- alecco 3y agoIt's in the Phi-1.5 technical paper. For phi-2 they bumped the number of tokens to 1.4 T and for sure most of it is generated, like previous models.
- IanCal 3y agoI might be missing it but I can't find where it says how the data was generated, it mostly refers back to the previous paper which started they used 3.5 I'd not be too surprised but I can't find anything in the technical report paper saying they're using 4 specifically.
- alecco 3y agoRead the first paper "Textbooks Are All You Need". > We annotate the quality of a small subset of these files (about 100k samples) using GPT-4: given a code snippet, the model is prompted to “determine its educational value for a student whose goal is to learn basic coding concepts”.
- IanCal 3y agoYes, they didn't use GPT-4 to generate data. They use GPT-3.5 to generate 1B tokens of synthetic data. They used GPT-4 to annotate data to train a classifier to filter human written code. The quote directly after yours: > We then use this annotated dataset to train a random forest classifier that predicts the quality of a file/sample using its output embedding from a pretrained codegen model as features. We note that unlike GPT-3.5, which we use extensively to generate synthetic content (discussed below), we use GPT-4 minimally only for annotations on the quality of a small subset of The Stack and StackOverflow samples. We thus view our usage of GPT-4 as merely a way to avoid tedious human-annotation efforts
- Der_Einzige 3y agoTraining lora's or other parameter efficient techniques to fine-tune LLMs can be done on a 3090 today for basically nothing.