3 ms·
Another recent (but not called out in this article) is the "Textbooks Are All You Need" paper [1]; the results seem to suggest that careful curation and curricu
by bluecoconut 3y ago
Another recent (but not called out in this article) is the "Textbooks Are All You Need" paper [1]; the results seem to suggest that careful curation and curriculums of training data can significantly improve model capabilities (when training domain specific, smaller models). Claiming a 10x smaller model can outperform competitors. (Eg. phi-1 vs. starcoder)
[1] https://arxiv.org/abs/2306.11644 https://arxiv.org/abs/2306.11644
- TaylorAlexander 3y agoI’ve still wondered if anyone has tried training with a large dataset of published books, like something from library genesis, or in the case of google using the full text from google books. There’s all this talk of finding quality text and I’ve not heard of text from print books being a major source beyond this textbooks paper?
- startupsfail 3y agoThat’s how OpenAI was (is) doing it. Books downloaded from the Internet is a part of the dataset as per GPT3 model card. Right to read.
- YetAnotherNick 3y agoTBH, it looks like metric manipulation to me. They have used GPT-3.5 to generate their data(and not use textbooks at all like the title suggests). And their dataset is very much like their benchmark data. While there was some filtering, but still it is very possible that lot of the benchmark questions were in training data. We likely wouldn't ever know how good the model is as it not only closed but they haven't provided access to anyone.
- bluecoconut 3y agoThey seemed to be pretty mindful of this contamination, and call out that they agressively pruned some training dataset and still observed strong performance. That said, I agree, I really want to try it out myself and see how it feels, and if the scores really translate to day-to-day capabilities. From section 5: In Figure 2.1, we see that training on CodeExercises leads to a substantial boost in the performance of the model on the HumanEval benchmark. To investigate this boost, we propose to prune the CodeExercises dataset by removing files that are “similar” to those in HumanEval. This process can be viewed as a “strong form” of data decontamination. We then retrain our model on such pruned data, and still observe strong performance on HumanEval. In particular, even after aggressively pruning more than 40% of the CodeExercises dataset (this even prunes files that are only vaguely similar to HumanEval, see Appendix C), the retrained phi-1 still outperforms StarCoder.
- YetAnotherNick 3y agoI read that but there is no good technique to rule out close duplicates. I know because I had tried to build one for my product. At best it relies on BLEU, embedding distance and other proxies which are far from ideal.
- dustypotato 3y agoIs there any clue on what architecture they used to create phi-1 ?