6 ms·
I’ve noticed a strong negative streak in the security community around LLMs. Lots of comments about how they’ll just generate more vulnerabilities, “junk code”,
by rpicard 1y ago
I’ve noticed a strong negative streak in the security community around LLMs. Lots of comments about how they’ll just generate more vulnerabilities, “junk code”, etc.
It seems very short sighted.
I think of it more like self driving cars. I expect the error rate to quickly become lower than humans.
Maybe in a couple of years we’ll consider it irresponsible not to write security and safety critical code with frontier LLMs.
- tptacek 1y agoThere are plenty of security people on the other side of this issue; they're just not making news, because the way you make news in security is by announcing vulnerabilities. By way of example, last I checked, Dave Aitel was at OpenAI.
- croes 1y ago[flagged]
- rpicard 1y agoFair! I’ve been surprised in some cases. I’m thinking specifically of a handful of conversations I was in or around during the Vegas cons. I might also be hyper sensitive to the cynicism. It tends to bug me more than it probably should.
- tptacek 1y agoA significant component of it is backlash/exhaustion from the volume of stories about it. I like LLMs, and I'm far past sick of stories about them.
- andrepd 1y agoLet's maybe cross that bridge when (more important, if!) we come to it then? We have no idea how LLMs are gonna evolve, but clearly now they are very much not ready for the job.
- bpt3 1y agoYou're talking about a theoretical problem in the future, while I assure you vibe coding and agent based coding is causing major issues today. Today, LLMs make development faster, not better. And I'd be willing to bet a lot of money they won't be significantly better than a competent human in the next decade, let alone the next couple years. See self-driving cars as an example that supports my position, not yours.
- philipp-gayret 1y agoWhat metric would you measure to determine whether a fully AI-based flow is better than a competent human engineer? And how much would you like to bet?
- bpt3 1y agoIn this context, fewer security vulnerabilities exist in a real world vibe coded application (not a demo or some sort of toy app) than one created by a subject matter expert without LLM agents. I'd be willing to bet 6 figures that doesn't happen in the next 2 years.
- anonzzzies 1y agoThe current models cannot be made to become better than humans who are good at their job. Many are not good at their job though and I think (see) we already crossed that. Certain outsourcing countries could have (not yet, but will have) millions of people without jobs as they won't be able to steer the LLMs to making anything usable as they never understood anything to begin with. For people here on HN I agree with you; not in the next 2 years or, if no-one invents another model than the transformer based model, not for any length of time until that happens.
- bpt3 1y agoAgreed. I think the parent poster meant it differently, but I think self driving cars are an excellent analogy. They've been "on the cusp" of widespread adoption for around 10 years now, but in reality they appear to have hit a local optimum and another major advance is needed in fundamental research to move them towards mainstream usage.
- kriops 1y ago> I think of it more like self driving cars. Analogous to the way I think of self-driving cars is the way I think of fusion: perpetually a few years away from a 'real' breakthrough. There is currently no reason to believe that LLMs cannot acquire the ability to write secure code in the most prevalent use cases. However, this is contingent upon the availability of appropriate tooling, likely a Rust-like compiler. Furthermore, there's no reason to think that LLMs will become useful tools for validating the security of applications at either the model or implementation level—though they can be useful for detecting quick wins.
- lxgr 1y agoHave you ever taken a Waymo? I wish fusion was as far along!
- rpicard 1y agoMy car can drive itself today.
- kriops 1y agoI, too, own a Tesla. And granted, analogous to the way we can achieve fusion today. Edit: Don’t get me wrong btw. I love autopilot. It’s just completely incapable of handling a large number of very common scenarios.
- mcv 1y agoYeah, there's a massive difference between a system that can handle a specific number of well-defined situations, and a system that can handle everything. I don't know what the current state of self-driving cars is. Do they already understand the difference between a plastic bag blowing onto the street, and a football rolling onto the street? Because that's a massive difference, and understanding that is surprisingly hard. And even if you program them to recognize the ball, what if it's a different toy?
- croes 1y agoIt’s the same problem as with self driving cars. Self driving cars maybe be better than the average driver but worse than the top drivers. For security code it’s the same.
- lxgr 1y agoRegardless of whether that comparison is valid: In a world where the average driver is average, that honestly doesn't sound too bad.
- croes 1y agoFor cars yes, but for security it would mean rolling your own crypto. There is a reason why the average programmer should use established libraries for such cases.
- kingstnap 1y agoFor now we train LLMs on next token prediction and Fill-in-the-middle for code. This exactly reflects in the experience of using them in that over time they produce more and more garbage. It's optimistic but maybe once we start training them on "remove the middle" instead it could help make code better.
- voidUpdate 1y agoYeah, it does sound a lot like self-driving cars. Everyone talks about how they're amazing and will do everything for you but you actually have to constantly hold their hand because they aren't as capable as they're made out to be
- xnorswap 1y agoI've been watching a twitch streamer vibe-code a game. Very quickly he went straight to, "Fuck it, the LLM can execute anything, anywhere, anytime, full YOLO". Part of that is his risk-appetite, but it's also partly because anything else is just really furstrating. Someone who doesn't themselves code isn't going to understand what they're being asked to allow or deny anyway. To the pure vibe-coder, who doesn't just not read the code, they couldn't read the code if they tried, there's no difference between "Can I execute grep -e foo */*.ts" and "Can I execute rm -rf /". Both are meaningless to them. How do you communicate real risk? Asking vibe-coders to understand the commands isn't going to cut it. So people just full allow all and pray. That's a security nightmare, it's back to a default-allow permissive environment that we haven't really seen in mass-use, general purpose internet connected devices since windows 98. The wider PC industry has got very good at UX to the point where most people don't need to worry themselves about how their computer works at all and still successfully hide most of the security trappings and keep it secure. Meanwhile the AI/LLM side is so rough it basically forces the layperson to open a huge hole they don't understand to make it work.
- tootubular 1y agoI know exactly the streamer you're referring to and this is the first time I've seen an overlap between these two worlds! I bet there are quite a few of us. Anyway, agreed on all accounts, watching someone like him has been really eye opening on how some people use these tools ... and it's not pretty.