4 ms·
You missed important context there. In particular, "These LLMs are being trained on data sets with bad results and bad code with no real way to tell the differe
by nanolith 2y ago
You missed important context there. In particular, "These LLMs are being trained on data sets with bad results and bad code with no real way to tell the difference."
A couple dozen bad SO articles can easily poison the results of thousands of examples of good OSS code. Code rarely has prose associated with it. SO articles have prose, so these articles will be disproportionately considered as the LLM is self-organizing.
So, it's not necessary that the LLM be trained primarily on bad SO articles for it to have a disproportionate impact on the results it generates from prose prompts.
- williamcotton 2y agoWhen I train a CNN the scale of errors is a very important characteristic of the training set so I empirically don’t understand this idea of “poisoning” with “a couple of dozen bad SO articles”. Do you have any sources related to this disproportionate impact?
- nanolith 2y ago"On the Dangers of Stochastic Parrots" (doi:10.1145/3442188.3445922) is a great introduction to this, even with its flaws. I also recommend Mitchell 2023 (doi:10.1073/pnas.2215907120) and Niven 2019 (arXiv:1907.07355) as good starting points. These don't directly address your question, but within the context of these papers, it's possible to see how the weighting of prose (which is a relation to prompts) and code can become skewed easily.