4 ms·
It seems like they should be able to “overweight” newer training data. But the risk is the newer training data is going to skew more towards AI slop than older
by matthewbauer 8mo ago
It seems like they should be able to “overweight” newer training data. But the risk is the newer training data is going to skew more towards AI slop than older training data.
- otabdeveloper4 8mo agoThere won't ever be newer training data. The OG data came from sites like Stackoverflow. These sites will stop existing once LLMs become better and easier to use. Game over.
- esclerofilo 8mo agoEvery time claude code runs tests or builds after a change, it's collecting training data.
- co_king_5 8mo ago[dead]
- esclerofilo 8mo agoI can't pretend to know how things work internally, but I would expect it to be involved in model updates.
- otabdeveloper4 8mo agoYou need human language programming-related questions to train on too, not just the code.
- 8note 8mo agothats what the related chats are for?
- otabdeveloper4 7mo agoAnd now you're training LLMs on LLM output. No, you need something like Stackoverflow. The crowdsourced ratings system that Stackoverflow has (had?) is the crucial part.