3 ms·
On the <|endoftext|>: GPT-2 and this model were trained by sampling fixed-length segments of text from a set of web pages. So if the sample happens to start nea
by AdamDKing 7y ago
On the <|endoftext|>: GPT-2 and this model were trained by sampling fixed-length segments of text from a set of web pages. So if the sample happens to start near the end of one page then it will fill in the rest of the length with the beginning of another page. The model learns to do the same. TalkToTransformer.com hides this by not showing what comes after the <|endoftext|> token.
- macawfish 7y agoThat explains why sometimes the talktotransformer samples are so short!