3 ms·
Hi, I don't have a comprehensive guide or anything like that, but I can quickly answer your questions: 1. Delimiting your text with <|endoftext|> lets the code
by serendipityrecs 6y ago
Hi, I don't have a comprehensive guide or anything like that, but I can quickly answer your questions:
1. Delimiting your text with <|endoftext|> lets the code know where the boundaries of your text chunks are. It makes sure that during training you're not inadvertently continuing past the end of a chunk of text into a totally unrelated one. However, it seems like there is a bug with this re how the nsheppard fork works: https://github.com/openai/gpt-2/issues/222 https://github.com/openai/gpt-2/issues/222
Even if you're not doing this, you shouldn't be getting gibberish, so you likely have another issue in your set up.
2. Finetuning just means training on top of a somewhat already trained model. GPT-2 is a language model, which means it's trying to learn (and generates from) a probability distribution over words (actually tokens). That probability distribution is going to be different depending on what corpus you're working with (wikipidia articles vs news vs reddit). GPT-2 (the models you download) is trained on a particular corpus but you as you fine-tune on your own corpus it's going to start pulling the learned probability distribution towards that of your dataset's.
I think the best way to play with this stuff is to run the code, like you're doing. When you get something that doesn't make sense, dig into the code/paper. There's also a lot of information in the readmes/issues for these repos.
- nonbirithm 6y agoThank you for the pointers, I'll rethink my strategy.