5 ms·
Split your training data into chucks of text that make sense. A random dataset example https://huggingface.co/datasets/imdb https://huggingface.co/datasets/imdb
by amrb 4y ago
Split your training data into chucks of text that make sense. A random dataset example https://huggingface.co/datasets/imdb https://huggingface.co/datasets/imdb
- meghan_rain 4y agoThanks, what does making sense mean? Be logically coherent (eg a paragraph of text in a document?) And does the training then create windows of ngrams on those chunks? Or what is the input/output? The reason I ask: If I had question/answer pairs, the question is the input, the answer is the output. What is the "output" when the input is just a (logically coherent) chunk of text?
- lxe 4y ago> What is the "output" when the input is just a (logically coherent) chunk of text? It probably won't change much if it's just a single sample. If you put in a large corpus of samples that repeat on the same theme, then the model will be "tuned" to repeat that theme. If you increase the number of epochs, you can overtrain it, meaning that it will just spit out the training data text.