4 ms·
From the article: "Our relative lack of skill at investigation becomes clear when we look at the accuracy rate of StackOverflow answers. For the amount of sass
by nanolith 2y ago
From the article:
"Our relative lack of skill at investigation becomes clear when we look at the accuracy rate of StackOverflow answers. For the amount of sass you see on that platform, you’d expect the programmers to at least be right. Except they aren’t. We have whole jokes about this too. Again, this is what was used to train the LLMs. Models trained on human data can’t outperform the base error rate in that data."
I think this is an important metric to consider. These LLMs are being trained on data sets with bad results and bad code with no real way to tell the difference. Only someone experienced in this field can read a SO article and appreciate the quality of the responses. An LLM is just summarizing bad information.
- Rinzler89 2y ago>These LLMs are being trained on data sets with bad results and bad code with no real way to tell the difference. No shit. I lost it when Google AI started recommending people to refill their blinker fluid, which is a meme that's sarcastically been parroted all over the internet for decades but Google's AI took it as serious gospel in their training dataset. The fact of the matter is these are just LLMs, not AI, they're only as good as their dataset. And since the entire internet is their dataset, along with all the human trolling, sarcasm and shitposting, and without any logic or intuition to distinguish what is sarcasm and what is not, their answers are gonna be questionably acurate a lot of the times when it comes to generic info that every average joe can have an opinion on. There's no actual intelligence there to call it an AI, it's just a summary of the entire internet, which is definitely useful for some task, but not at replacing human tasks. I saw some people on HN defended Google AI's shitty answers as being very useful for humor and not for real answers, but then that would be the world's most expensive joke machine for a company who's main business is search and thrived as the go-to for web answers. You don't need to delegate 3% of the world's energy, or whatever it is, towards generating only wrong answers that sound funny.
- ToucanLoucan 2y agoAt this point I adamantly refuse to call it AI one single more time, because it is not, it's LLM's, it's generative models, it's pretrained transformers. AI as a term belies an intelligence in the product in your average person (and many tech people it seems) that is simply not. Fucking. There. LLMs do not understand, at all, what they are saying. They are simply generating words one at a time to minimize "error," whatever that means in the context of the model.
- jajko 2y agoSome malicious folks concerned about stability of their work may decide to poison training data... not an easy task for sure, but not entirely impossible in their niche area of expertise. Like literally all powerful tools in the past, ai abuse will come from various directions for various unexpected reasons.
- sokoloff 2y ago“Models trained on human data can’t outperform the base error rate in that data.” That doesn’t seem inherently/unavoidably true. Surely a human could read a bunch of SO answers, try them, and thereby outperform the base error rate of the answers. I don’t see it as impossible for an AI model to exist that does the same.
- ToucanLoucan 2y agoBecause a human can comprehend, and an LLM cannot.
- jqpabc123 2y agoI don’t see it as impossible for an AI model to exist that does the same. The important step is creating an AI model with enough cognition to recognize it needs to do the same before spouting nonsense.
- nanolith 2y agoSuch a model has not yet been invented. Certainly, a generative model such as a large language model is unlikely to gain the ability to reason about these answers. LLMs are a step on the path toward better human/computer interaction, but they are just a step. Like most new things in the AI field, they have been over-hyped and have moved the market in a giant game of tech FOMO.
- jqpabc123 2y agoa giant game of tech FOMO. For me, the most distressing part is how easily tech people drink the FOMO and then go to work promoting it. This happens over and over again.
- jqpabc123 2y agoAn LLM is just summarizing bad information. In other words, there is no real *intelligence* involved at all. *Intelligence* has yet to be accurately defined --- and there is no reason to believe it is a statistical function.
- ben_w 2y agoWhen I was a kid — and not just as a kid, this was still the case 10 years ago — summarising text was considered "AI-complete" (also known as "AI-hard"). In fact, at time of writing, "natural language understanding" is still on the Wikipedia page for AI-complete. > Intelligence has yet to be accurately defined --- and there is no reason to believe it is a statistical function Other than the work of Rev. Bayes, Prof. Schrödinger. Plus, now I think about it, saying "randomness can't lead to intelligence" is the talking point for all the people who insist Darwin is wrong and there has to be an intelligent designer.
- renewedrebecca 2y agoWhat LLMs do and what Evolution does are not the same thing. Evolution is a process where small changes over a long time create new, more complex organisms. LLMs are not that. LLMs are not evolving themselves, giving themselves new abilities. Just adding more variables to a long list of variables is not the same thing.
- jqpabc123 2y agosaying "randomness can't lead to intelligence" is the talking point Given enough time, a monkey randomly typing on a keyboard could hypothetically produce Shakespeare. But talking points aside, this does not make the monkey more "intelligent" nor does it mean that "randomness" is a practical methodology for achieving specific results on a human timescale. Particularly when the desired result is as complex and poorly understood as "intelligence".
- ben_w 2y agoI'm describing evolution itself as the intelligence here, not the monkey. And the million typewriters taking ages is why Creationists argue it can't work (having not understood half of it). > But talking points aside, this does not make the monkey more "intelligent" nor does it mean that "randomness" is a practical methodology for achieving specific results on a human timescale. Not by itself. Evolution is more than just random, it's a specific filter on top of randomness, just like all the other things. And computers roll the dice very quickly compared to reproduction, which is why neither you nor any other human can win a chess game against the best computer any more — it did it on a human timescale.
- paradox242 2y agoExcept it's not even summarizing, it's generating new and creative ways to be wrong.
- williamcotton 2y agoAn LLM is just summarizing bad information. You cannot make this claim without evidence that the LLM primarily used poor quality SO articles. You don’t include an assessment of all of the other sources. Other sources could include properly functioning open source code, for example.
- nanolith 2y agoYou missed important context there. In particular, "These LLMs are being trained on data sets with bad results and bad code with no real way to tell the difference." A couple dozen bad SO articles can easily poison the results of thousands of examples of good OSS code. Code rarely has prose associated with it. SO articles have prose, so these articles will be disproportionately considered as the LLM is self-organizing. So, it's not necessary that the LLM be trained primarily on bad SO articles for it to have a disproportionate impact on the results it generates from prose prompts.
- williamcotton 2y agoWhen I train a CNN the scale of errors is a very important characteristic of the training set so I empirically don’t understand this idea of “poisoning” with “a couple of dozen bad SO articles”. Do you have any sources related to this disproportionate impact?
- nanolith 2y ago"On the Dangers of Stochastic Parrots" (doi:10.1145/3442188.3445922) is a great introduction to this, even with its flaws. I also recommend Mitchell 2023 (doi:10.1073/pnas.2215907120) and Niven 2019 (arXiv:1907.07355) as good starting points. These don't directly address your question, but within the context of these papers, it's possible to see how the weighting of prose (which is a relation to prompts) and code can become skewed easily.