8 ms·
We can chalk this up as another example of over-exhuberance by what folks believe LLMs can accomplish vs. what they actually are. LLM-based “AI” is able to use
by gortok 2mo ago
We can chalk this up as another example of over-exhuberance by what folks believe LLMs can accomplish vs. what they actually are.
LLM-based “AI” is able to use its vast corpus of inputs and calculate the most statistically likely output in a given situation. It is probabilistic, and when you are dealing with probabilities in a situation where certainties, not probabilities, matter, you’re going to get dinged on credibility massively when your LLM-based “AI” gets the probabilities wrong at best, or in this case, claims a line of code generates a vulnerability when it is, in fact, a code comment.
LLMs are text-prediction engines. They are not Artificial Intelligence, and shouldn’t not be treated in any form or fashion as if they possess intelligence. What bothers me about this entire situation is that presumably the folks that relied on the LLM-based “AI” to generate these vulnerabilities knew (or should have known) enough about their tool to know this would happen, but did not.
Now, we all pay the consequence, to the tune of hundreds of thousands if not millions of dollars of wasted productivity from teams that have to deal with the resulting fall-out of this usage of “AI”.
A human must verify everything an LLM presents as fact. Everything. If you don’t, we all pay the price. LLMs do not remove the onus of responsibility on the human being, if anything they amplify it because LLMs can generate lots more output more quickly that needs to be verified than humans can.
- geraneum 2mo agoUnfortunately people sometimes get defensive against this take. But I think treating the LLM as you described can make you a better LLM user and help get better output. It helps understand the failure modes better, and moderate one’s reliance on them. Just like how we should do for every tool we work with.
- gr_norm 2mo agoYes, I've found that reminding yourself of how they actually work helps keep you on guard against LLM-patterned mistakes. Especially things like carefully considering what parts of the current task likely fall outside the distribution of corpus + RL data (as much as that can be guessed).
- palmotea 2mo ago> ...Unfortunately people sometimes get defensive against this take. But I think treating the LLM as you described can make you a better LLM user and help get better output. It helps understand the failure modes better, and moderate one’s reliance on them. Just like how we should do for every tool we work with. B...b...but the Anthropic trainer said we'd get the best results if we don't think of it as a tool, but instead give it a name and think of it as our brilliant coworker! Why should I trust you, internet rando over a stormtrooper-level salesman? /s
- red75prime 2mo agoApophatic intelligence? "We don't know what intelligence is, but LLMs with CoT are certainly not it despite being Turing-complete." Watching for unexpected failure modes is surely worth it.
- gortok 2mo agoTuring-completeness is a necessary pre-requisite for being able to fulfill the requirements of a Turing machine, nothing more. In the same way that cell division is a necessary condition for life, but cell division does not mean a given life form itself is sentient. Intelligent life-forms can generate probabilistic outputs based on inputs, but being able to generate probabilistic outputs based on inputs is not what makes us intelligent.
- estearum 2mo ago> but being able to generate probabilistic outputs based on inputs is not what makes us intelligent. ??? Of course it is. The brain is mechanically not capable of doing anything other than that. Do you believe the brain is something other than a bundle of probabilistic physical interactions? Or are brains not the source of what we call intelligence?
- nullsanity 2mo ago[dead]
- pessimizer 2mo agoYours is a controversial view. It is lazy and selfish to try to get other people to explain their case that it is not exclusively that, when saying that it is exclusively that is the weaker case, and you back it up with nothing but a snarky proclamation. Are newly born babies reacting due to statistical probabilities that they have derived, or are they using something other than their brains?
- estearum 2mo ago
- tsunamifury 2mo ago[flagged]
- elmer2 2mo agoMany people with no skills are taking advantage of the LLM craze to artificially inflate their own value. I see it every day on LinkedIn. People that previously have barely any experience in tech, now being hired in AI startups because they are good bullshitters.
- prh8 2mo agoCountless directors and managers are now cosplaying as engineers. I've seen so many myself and that's just my tiny slice of this engineering world
- 1-6 2mo agoEngineers cosplay as physicists and mathematicians every day. What's your point? Think of it pragmatically. If they can do the job they can do the role.
- ruszki 2mo ago> If they can do the job they can do the role. Obviously. Can they do the job? Because right now, government decisions are based on AI generated code, which was verified by nobody who can do that. So the cost of an unsatisfactory answer is quite high.
- Blackthorn 2mo agoIt is pretty funny to see the shoe on the other foot, since it's usually software engineers with unearned arrogance about other fields.
- ryan_n 2mo agoWhat fields do you see devs think they know about? I’ve never personally seen this with other devs I work with but obviously small sample size…
- mathisfun123 2mo ago
- bwfan123 2mo ago> Now, we all pay the consequence, to the tune of hundreds of thousands if not millions of dollars of wasted productivity from teams that have to deal with the resulting fall-out of this usage of “AI”. Brandolini's principle in action. It takes 10 times more energy to refute BS than to generate it. A related analogy to computing: it is easy to generate propositions, but hard to test if a given proposition is satisfiable or not, which curiously ties to P vs NP.
- Sohcahtoa82 2mo ago> Brandolini's principle I much prefer the alternative name: the Bullshit Asymmetry Principle.
- Jblx2 2mo agoSeems like most of it is covered by: Entropy increases.
- ivan_gammel 2mo agoYou are right with the analysis, but wrong with the conclusions. Yes, LLM „thinking process“ is kinda non-deterministic in a sense that it does not follow logical reasoning and will not produce logically correct results in 100% cases. It has an error margin. However, error margins are in the center of any engineering discipline. We cannot produce things measured with 100% accuracy. This is accepted fact. The focus is always not on eliminating errors, but on reducing them to acceptable minimum. With LLMs we should not expect an ideal logical thinker, but a process that may error sometimes, and we must design quality controls instead that push LLM outputs within acceptable margins. And it can work.
- kentm 2mo agoYes but the key here is doing proper risk assessment. "What is the consequence if the LLM gets this wrong?" "How do we verify the output?" "What are the legal ramifications for using the LLM in this way?" "Who is responsible when the LLM fails?" "Whats the expected accuracy here?" etc. In the current AI mania, there's a lot of due diligence simply being ignored. Plenty of "Well humans make mistakes too!" going on here on HN too.
- ivan_gammel 2mo agoThe due diligence not being done is people putting cats in microwaves. It‘s not the dangerous part. The real danger is risk assessments coming to wrong conclusions, because it is still terra incognita. Talented engineers were in this situation before, doing mistakes with cars, airplanes, buildings etc.
- kentm 2mo agoNo, I'm sorry but I think thats a cop out. The fact that LLM are stochastic and can give incorrect answers is not particularly difficult to comprehend, and the risks that fall out of that are reasonably understandable. The issue is entirely down to bad choices by the people driving LLMs, because they are engaging with what they wish LLMs do instead of what they actually do.
- budsniffer952 2mo ago[flagged]
- deleted 2mo ago[deleted]
- sedawkgrep 2mo ago> We are not going back, period. I didn't get this at all from the parent. They're simply stating that LLMs aren't entirely trustworthy, and that the responsibility is ultimately ours, not the LLM's.
- budsniffer952 2mo agoOkay.
- gbnwl 2mo agoEvery day I wake up and open HN. “LLM has made legitimate mathematical discoveries” —> Wow the rate of progress is amazing. Highly upvoted. “LLM does something not good” -> Does everyone else not realize LLMs are just dumb next token predictors? Highly upvoted. So tired of this discourse and this site.
- apples_oranges 2mo agoWould be nice to get high karma commenter votes count only ..
- Jensson 2mo agoThe rate of progress can be high and they can also be dumb next token predictors. Not sure why that is hard to understand. These models can do a lot of things but they also can't do a lot of things. In order to use these models effectively you have to understand that they are next token predictors and how that allows it to do what they do.
- gbnwl 2mo agoAre they useful or not? Will they continue changing the world or not? People who choose one way or the other for describing them typically fall on one side or the other in these questions imo. What do you think? Will these next token predictors change the world or not?
- Jensson 2mo agoThey are useful. They will continue to change the world. They are still next token predictors with all the problems that comes with that. For them to change the world you have to work with them as next token predictors. Ensure that the next token predictor has enough prediction paths to solve the problems you want and so on. Since when they don't they fail spectacularly. These big companies will continue to add new skills to them, so they will continue to get more useful.
- deleted 2mo ago
- SubiculumCode 2mo agoWe can chalk this up as another example of over-exhuberance by what folks believe humans can accomplish vs. what they actually are. Flesh-based “brain” is able to use its vast corpus of inputs and calculate the most statistically likely output in a given situation. It is probabilistic, and when you are dealing with probabilities in a situation where certainties, not probabilities, matter, you’re going to get dinged on credibility massively when your flesh-based brain gets the probabilities wrong at best, or in this case, claims a line of code generates a vulnerability when it is, in fact, a code comment. Humans are prediction engines. They are not Pure Intelligence, and shouldn’t not be treated in any form or fashion as if they possess pure intelligence. What bothers me about this entire situation is that presumably the folks that have relied on the flesh-based “brains” to generate these vulnerabilities knew (or should have known) enough about their "tool" to know this would happen, but did not: To err is to be human. Now, we all pay the consequence, to the tune of hundreds of thousands if not millions of dollars of wasted productivity from teams that have to deal with the resulting fall-out of this over reliance on fallible “brains". A human must verify everything another human presents as fact. Everything. If you don’t, we all pay the price. Using a human does not remove the onus of responsibility on the human being in charge, if anything they amplify it because humans work for peanuts in some countries, and can generate lots more output more quickly that needs to be verified by the humans in charge.
- adjfasn47573 2mo ago> A human must verify everything an LLM presents as fact. Everything. I've thought about this for quite some time now. No. A human doesn't need to verify everything. And the argument is really simple: stochastic. Think of self-driving cars: We can show today - based on evidence and real data - that self-driving cars are safer than human drivers. That's a fact and the consequences are clear, more self-driving cars, less human-driven cars, less accidents, less hurt people, less dead people. Are the cars 100% safe and NEVER make a mistake? No. But they don't need to. Nothing is ever 100% (in the real world). Now back to AI for software creation. "Review is the bottleneck because EVERYTHING must be judged by a human." No. It doesn't. We just need to build AI review systems, that will do reviews better than (or at least as good as) humans. The human review quality bar is far below 100%. Far far far. If we can show (likely in the next 12-24 months I think) that AI review quality is consistently above the human review quality - again, based on evidence, based on real data - then that's it, then there's no good reason to have humans review the code. Yes, there will be another layer in the system, another level of abstraction that will/must end at the human boundary.
- gspr 2mo agoThis reduction of everything to stochasticity is silly. Or, to put it differently: Do you accept a value with some error appearing in your bank account on salary day? We have plenty of systems where complete accuracy is the only acceptable thing. Computers are great for such things. Until we all get caught up in a way of delusion and start writing those systems as natural prose passed through an improperly understood stochastic machine.
- adjfasn47573 2mo agoI've never said or implied what you're arguing against right now. An LLM building software benefits from using a formal language, formal tests or having a formal specification in a similar way we do. Everything you can formalize into something, so you have certainty about stuff, is always a benefit. I wasn't comparing exact science with LLMs. I was comparing current messy review processes of code performed by humans with future messy review processes of code likely performed by LLMs. If you find a way to convert code reviews into a fully formalized process, then this is clearly the winner. If the LLMs find a way to do that, same. If you find a way to formalize the process of driving through any street in any situation, then this is clearly the winner. Until then, driving stays messy and stochastic.
- aaroninsf 2mo agoThis is a conflation of issues, predicated on false understanding of what LLMs are. This line of critique is pernicious because it is both technically correct, as description, and profoundly misleading. Saying that outputs are a product of inputs is not interesting and to the point it is not explanatory. What is interesting, is how they do what they do. What is the "statistically likely* next token? To answer that you can do exactly one thing, run the LLM. That's because what they are doing is interesting and not reducible. What is more interesting is that in order to do what they do, given the architectures we apply and the training strategies we use and the harnesses we situate them in, LLM are recapitulating in their deep layers strategies observed in the animal brain. This is still suggestive, interpretibility is nascent: but it is also more than a little interesting. In some respects, for cognitive scientists interested in the manner in which mind merges from computational substrates, it is profoundly interesting. One can incorporate this, and, still be viciously critical of bother the success and failure of LLM in the applications we have put them to, and of how we (as individuals and as institutions such as corporations) are integrating them into our work. There is a lot to criticize! But criticism can be taken more seriously when it is not obscured by misunderstanding or misrepresentation (intentional, or not) of what LLM are and why they are not remotely "parrots" in the pejorative sense. The technology, as technology, at the scale we are architecting it, is doing things we did not imagine would be witnessed in our lifetime, if ever. Dismissing that and denying it because of the career, industry, society, and civilization challenges that technology brings are existential, is bad argumentation or bad faith. Both can be true at once.
- a3mrouter1 2mo ago[flagged]
- chrisjj 2mo ago> We can chalk this up as another example of over-exhuberance by what folks believe LLMs can accomplish vs. what they actually are. I see no credible corroboration. More likely its folks having no more care for what they are doing than the bots themselves. > Now, we all pay the consequence, to the tune of hundreds of thousands if not millions of dollars of wasted productivity from teams that have to deal with the resulting fall-out of this usage of “AI”. People said the same about email spam ... until they engaged spam filters. CVE report slop is simply spam. Complaints are better directed at the filters, not the filtered.
- polotics 2mo agoWow you just got us a complete nostalgia moment to the good old times when the computer who always beats us at chess became `not artificial intelligence`...
- treszkai 2mo ago> LLMs are text-prediction engines. They are not Artificial Intelligence, and shouldn’t not be treated in any form or fashion as if they possess intelligence. I agree that humans must verify LLM-produced facts, but strongly disagree with these kinds of "stochastic parrot therefore dumb" arguments. Yes, an LLM is a "stochastic parrot". No, that doesn't imply that it is dumb. Enough to look at how Terence Tao asks ChatGPT to help him understand a solution that nobody had ever discussed before [1], or how a random guy asks ChatGPT in a handful of words to disprove a 30-year-old conjecture, with zero technical input [2]. If your parrot in a birdcage with internet access can finish the sentence, "The counterexample to the Dinitz–Garg–Goemans conjecture is...", then it's a pretty smart parrot, by all reasonable definitions of "smart". Just because someone bottled up the formula into matrix multiplications and added some random sampling to the outcome, that doesn't take away from the fact that the parrot said provably correct statements that the biggest experts in the field couldn't imagine. And no, I'm not implying that the LLMs are correct all the time, or that their intelligence and reasoning works in any way like ours. [1]: https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed56 https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed... [2]: https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de063 https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de0...
- Jensson 2mo agoThe best way to describe the LLM intelligence is "an expert system that works the way people thought expert systems would work". You can encode a massive amount of skills into an LLM, and then the LLM uses those to navigate problems. But the LLM is still dumb where those skills doesn't have good coverage, since unlike the expert systems it maps fuzzily to its skills, and they are tuned to produce results over rejecting the request when its unclear if coverage is good. As long as that is true you have to treat them as dumb even if they sometimes produce brilliant results.
- gortok 2mo agoI’m not looking to argue about your position, but I think the inciting incident in the OP shows evidence that LLMs are “dumb” even when they have the sum total of written language and code at their training disposal. In this case I would expect, for your supposition to be true, that an LLM would not mistake a code comment for an actual vulnerability.
- vouwfietsman 2mo agoDoes a dog possess intelligence? Does a bird? Does a cricket? An amoeba? I hate AI slop as much as the next guy but the amount of tribalism over AI is taking near-religious forms. Nobody knows what intelligence is, therefore we don't know what does or does not possess it, therefore we don't know whether LLMs currently, or in the future, possess it. Yes, LLMs can be stupid, guess what: so can I. That doesn't really change the argument at all. I feel like I'm on a deja-vu from when DALL-E was released and everybody was fighting over whether AI can be creative yes or no. Same story, different words. Intelligence, creativity: we have no idea what these words mean, and AI is helping us understand them better. That alone is an achievement of epic proportions. I am not joking here. Any computer scientist before 2015 would be absolutely blown away by what you can now do for 10 cents and an API call, yet somehow because of the tech-bro-iness of it all we get a tribal war over what is plainly visible in front of us: LLMs are uncomfortably close to what we thought intelligent machines would look like
- Jensson 2mo ago"Dumb next token predictor" keeps popping up since that is the core way they work. Since they aren't logic engines but prediction engines they will always return a result regardless what you ask it. Some predictions might be the tokens "I don't know", but that is based on the model mapping your text to those tokens by having seen many similar "I don't know" responses to such contexts, it didn't do any introspective logic to produce that "I don't know", and its possible it actually does know if it followed another branch there so "I don't know" is often not even true. If they had an introspective part that stops the prediction when its too unreliable it would no longer just be token prediction engine, and I believe we need such a part for them to become what I call smart. I don't think LLM will ever stop being dumb without such an introspective part to them. And no, that introspective part is not a part of the token predictor. At least not in us humans, the feeling of certainty we have is not a prediction, it is bundled with our thoughts, so we get both "answer is a bear" and "certainty is low", we don't get just one of those as a "prediction". Will LLM become smart as humans with such an introspective part? I don't know, but I think they will never become as smart as humans without one. Note: The certainty score has to be per conclusion or response, not per token. You can't evaluate a responses validity by aggregating the weight of each token. Meaning its a logic engine, not token engine, that evaluates the certainty of a statement being correct or not instead of a token being correct or not. That is the level human thinking works at and seems to be dramatically more efficient.
- bigbuppo 2mo ago> A human must verify everything an LLM presents as fact. Everything. If you don’t, we all pay the price. LLMs do not remove the onus of responsibility on the human being, if anything they amplify it because LLMs can generate lots more output more quickly that needs to be verified than humans can. The sort of person that's going to offload their thinking to AI is the exact sort of person that is not going to verify anything because they've already offloaded their thinking to AI.
- gortok 2mo agoFor a long time I was anti-licensure in tech; now with the bar being lowered to next to nothing, it seems as if licensure is more important than ever — not to protect this trade (though it will do that, and that is a benefit), but because the sheer amount of irresponsibility in the usage of LLMs and “AI” in general begs for licensure and adoption of a regulatory body for software in general.
- rafterydj 2mo agoThis is unfortunately a feeling I share. It wasn't until LLMs have become nearly ubiquitous at this point, and there has been zero realistic technological response to the dangers they present. Not to mention I suspect there may be some psychological element to being exposed to interactions with AI models and their nonsense for hours a day. Not all of it is nonsense....but you won't ever know for sure.
- 27183 2mo ago100%. I don't think "AI" has changed anything wrt to the responsibility of the tech industry in general. Tech has always been pretty much devoid of ethics or a sense of responsibility at the executive level (and therefore "management" more broadly). But before LLMs it was easier for ethical engineers to surreptitiously steer things towards responsible implementations. I suspect also competence correlates pretty strongly with intellectual honesty--narcissists and other personality defectives generally aren't terribly capable of the kind of introspection necessary to learn deeply. With the deluge of LLM slop, it's much harder for even a responsible manager to differentiate the competent from the incompetent, and now the dishonest, unethical folks have more leverage. So now the problem is highly visible. There's only one solution I can think of, and that's to hold individuals (not corporations) responsible for professional malfeasance.