4 ms·
Around of quarter of GPT's training data was Reddit so in some sense it's already a Reddit response generating API.
by CSMastermind 4y ago
Around of quarter of GPT's training data was Reddit so in some sense it's already a Reddit response generating API.
- Cipater 4y agoThis can't be right. Any place I can read about this?
- CSMastermind 4y agohttps://arxiv.org/pdf/2005.14165.pdf https://arxiv.org/pdf/2005.14165.pdf WebText and WebText2 referenced in their papers are corpuses based on Reddit submissions which had a 22% weight in their training model. https://openwebtext2.readthedocs.io/en/latest/ https://openwebtext2.readthedocs.io/en/latest/ This is larger than Wikipedia (3% weight) or either of their two book corpuses (8% each). The only other data included was a filtered set from Common Crawl (weighted 60%). --- I was imprecise with my language before but hopefully that at least provides some clarity.