3 ms·
Indeed but this is zero-shot performance. Fine-tuning for a task should get you pretty good results. I'm interested in seeing the results of an Alpaca method
by binarymax 4y ago
Indeed but this is zero-shot performance. Fine-tuning for a task should get you pretty good results. I'm interested in seeing the results of an Alpaca method against this Cerebras 13B model.
- MacsHeadroom 4y ago>I'm interested in seeing the results of an Alpaca method You're talking apples to oranges. The "Alpaca method" is a dataset generation method. Nothing about Alpaca's training method is novel, interesting, or efficient. Alpaca used the same standard training method everyone else uses, A100 clusters. If you mean LoRA/PEFT training which people used to replicate Alpaca then that is also apples to oranges because LoRA/PEFT is a finetuning method not a pre-training method.
- deleted 4y ago[deleted]
- UncleEntity 4y agoOne could take the alpaca dataset and fine tune using the LoRA/PEFT method and compare to the Stanford alpaca fine tuned llama model. Presumably…
- Vetch 4y agoBase model performance is what's most important and also impacts fine-tuning quality. Practically, a model that's good out of the box with minimal fine-tuning is also useful to more people. Since they focused on being training compute optimal for some budget, expect their models to lag behind Llama overall. Their 6.7B version should lag behind GPT-J, assuming 20 tokens per parameter. The Pythia models are also worth checking out, they might be better than or matched to CerebrasGPTs at each size (although they warn it is not intended for deployment). Conclusion: the landscape of top open models remains unchanged.
- boomerang90 4y agoI agree fine-tuning for task will give better results. Cerebras actually showed some research recently on this front. Sparse pre-training and dense fine-tuning (https://arxiv.org/abs/2303.10464 https://arxiv.org/abs/2303.10464). You can recover the accuracy of sparse pre-trained models with dense fine-tuning and reduce FLOPs of the end-to-end pipeline by 2.5x compared to dense.