3 ms·
> But I assume the data cleaning process removes such content before pretraining? ;) I didn't check what you're referring to but yes, the major providers likel
by throwaway314155 1y ago
> But I assume the data cleaning process removes such content before pretraining? ;)
I didn't check what you're referring to but yes, the major providers likely have state of the art classifiers for censoring and filtering such content.
And when that doesn't work, they can RLHF the behavior from occurring.
You're trying to make some claim about garbage in/garbage out, but if there's even a tiny moat - it's in the filtering of these datasets and the purchasing of licenses to use other larger sources of data that (unlike Common Crawl) _aren't_ freely available for competition and open source movements to use.
- jedimastert 1y ago> purchasing of licenses to use other larger sources of data https://www.npr.org/2025/09/05/g-s1-87367/anthropic-authors-settlement-pirated-chatbot-training-material https://www.npr.org/2025/09/05/g-s1-87367/anthropic-authors-...