3 ms·
They trained on https://huggingface.co/datasets/bigcode/the-stack-dedup https://huggingface.co/datasets/bigcode/the-stack-dedup which is a massive curated datas
by runnerup 3y ago
They trained on https://huggingface.co/datasets/bigcode/the-stack-dedup https://huggingface.co/datasets/bigcode/the-stack-dedup which is a massive curated dataset accumulated from GitHub. Details are here: https://www.bigcode-project.org/docs/about/the-stack/ https://www.bigcode-project.org/docs/about/the-stack/
Many of the most-represented "languages" on GitHub are actually things like JSON, XML, HTML, CSV, text, markdown, YAML, and SVG.
More details from them here: https://blog.replit.com/llm-training https://blog.replit.com/llm-training