4 ms·
There is one already: https://arxiv.org/abs/2305.07759 https://arxiv.org/abs/2305.07759 https://huggingface.co/datasets/roneneldan/TinyStories https://huggingfa
by pxagntuvzt 4mo ago
There is one already:
https://arxiv.org/abs/2305.07759 https://arxiv.org/abs/2305.07759
https://huggingface.co/datasets/roneneldan/TinyStories https://huggingface.co/datasets/roneneldan/TinyStories
6.5GB of tiny stories, as requested. ;)
- sixtyj 4mo agoTexts in Gutenberg have 20GB, and full Wikipedia (English texts) have 80-110GB. So to LLM-generate 6.5GB of tiny stories is quite a permutation in action :)
- Lerc 4mo agoMy comment was, in-fact, a subtle reference to this. The best opening I got from my own TinyStories trained model was. Once upon a time, in a small town, there was a large town. Which I just love as an evocative idea.
- aesthesia 4mo agoSimpleStories is a more diverse version: https://huggingface.co/datasets/SimpleStories/SimpleStories https://huggingface.co/datasets/SimpleStories/SimpleStories