4 ms·
Obviously I can only speculate since I neither have access to their dataset nor interest in paying for API access, but crawling and dataset cleaning have gotten
by strangecasts 1y ago
Obviously I can only speculate since I neither have access to their dataset nor interest in paying for API access, but crawling and dataset cleaning have gotten much better since the GPT-2 days, especially after Microsoft's PHI models [1] demonstrated how much dataset construction matters for parameter efficiency and toxicity. Having some basic content filtering is a pretty established part of data cleanup -- e.g. the fastText toxicity classifiers in the Dolma pipeline [2] -- which obviously still leaves in bad data, but certainly won't leave in the entirety of /b/
If shoddy data collection was the problem, we should expect the model to do much worse on overall leaderboards like [3], which require models to answer questions without sudden detours into Holocaust denialism. A change to the system prompt is more consistent with this, and as an added benefit, only requires one person to be completely out of their gourd.
[1] https://www.microsoft.com/en-us/research/publication/textbooks-are-all-you-need-ii-phi-1-5-technical-report/ https://www.microsoft.com/en-us/research/publication/textboo...
[2] https://arxiv.org/pdf/2402.00159 https://arxiv.org/pdf/2402.00159
[2] https://livebench.ai/#/ https://livebench.ai/#/