3 ms·
A related question: are the 40+ [1] millions Full Text Books used ? OpenAI is using book1 (BookCorpus ?) and book2 sources. By the number of tokens, this seems
by CrypticShift 4y ago
A related question: are the 40+ [1] millions Full Text Books used ?
OpenAI is using book1 (BookCorpus ?) and book2 sources. By the number of tokens, this seems less than a million book in total.
[1] https://www.blog.google/products/search/15-years-google-books/ https://www.blog.google/products/search/15-years-google-book...
- jll29 4y agoIt's "BooksCorpus" (with an 's'), a 800M word dataset described in Zhu et al. (2015) IEEE ICCV, and also available on AWS at: https://aws.amazon.com/marketplace/pp/prodview-d3ghxqzkitn6y https://aws.amazon.com/marketplace/pp/prodview-d3ghxqzkitn6y The Google BERT paper (Devlin et al., 2018) also references it: https://aclanthology.org/N19-1423/ https://aclanthology.org/N19-1423/ Privacy questions aside (as important as they of course are), it's very important to know what a model was trained on exactly: if Wikipedia was used in the training set, you can't use questions from Wikipedia to test it (as that would be cheating) - test data must be as "unseen" as a good exam.
- touringa 4y agoActually, it's 'BookCorpus'. OpenAI spelt it wrong in their GPT-1 paper. It has also been analyzed here and here: https://arxiv.org/abs/2105.05241 https://arxiv.org/abs/2105.05241 https://lifearchitect.ai/whats-in-my-ai/ https://lifearchitect.ai/whats-in-my-ai/