3 ms·
> In the real world, it does feel likely that we’re going to hit some sort of a ceiling on the number of useful bugs, and probably we’ll hit it soon. This does
by mbroshi 2mo ago
> In the real world, it does feel likely that we’re going to hit some sort of a ceiling on the number of useful bugs, and probably we’ll hit it soon.
This doesn't resonate with me. I see companies adding more sloppily written features with AI. I see more bugs in the software I use, not less. While it's plausible that software is getting both buggier and more secure, I suspect those two move in the same direction not opposite.
My guess is that we're getting better at finding _existing_ security issues with AI (and thus fixing those issues), but simultaneously adding more insecure surface areas _at a faster rate_.
- aleksandrm 2mo agoI don't know, my colleague refuses to use AI and I've been seeing more bugs from their side, while reducing bugs on my side with the help of AI. That said if companies want to "ship ship ship fast", then yes even AI can produce bugs or regressions if not carefully reviewed by the human.
- bossyTeacher 2mo ago> I've been seeing more bugs from their side, while reducing bugs on my side with the help of AI. You should question your ability to see any bugs on YOUR side.
- Ancapistani 2mo agoI don't have any colleagues like that anymore, but even as far back as the last half of 2025 I was seeing that automated AI review was becoming effective enough that I considered it essential to any project where security was a serious concern. These days we're generating multiple times more code than we were writing before. That means a similar multiple of opportunities for bugs to be introduced - so the ability to automate security review is more impactful in proportion to that.
- Forgeties79 2mo agoHave you asked your colleagues what it’s like to deal with your code?
- boc 2mo agoIt's not 2025 anymore my friend.
- Forgeties79 2mo agoYet apparently people still dump unvetted LLM outputs onto their colleagues and expect them to thank them for the privilege. So it’s worth asking them what the consensus is of their work to find out if it’s the case.
- leptons 2mo agoThat's the old way. The current way is nobody reviews anything. The LLM reviews it all and nobody even reads it, they just click approve, and merge blindly. I wish I were kidding.
- mlrtime 2mo agoNot where I work. We have both review PRs.
- DANmode 2mo ago> nobody Speak for your workplace only. Most places still consider a PR or commit yours, if your name is attached. Act accordingly.
- leptons 2mo ago>Act accordingly I've been looking for a new job where people aren't so reckless... the AI psychosis is heavy at my current job.
- kelnos 2mo ago> while reducing bugs on my side with the help of AI. How do you know? You might be adding (latent) bugs every time your LLM fixes one for you.
- remus 2mo agoSame as usual: add tests. Over time the test suite becomes the spec that describes how the software should function. As the test suite grows you squeeze out room for undefined behaviour and bugs.
- coldtea 2mo ago>Same as usual: add tests. Over time the test suite becomes the spec that describes how the software should function. That covers functionality - it doesn't catch the kind of bugs talked about in TFA.
- tonyedgecombe 2mo agoI remember decades ago IBM published some research that said that every two bug fixes resulted in a new bug.
- budman1 2mo agoAs long as the first derivative is negative, you can finish with a usable product. If the first derivative is positive, adding people does not help.
- tonyedgecombe 2mo agoYes, it just means your bug queue is twice as long as you think it is. https://en.wikipedia.org/wiki/1/2_%2B_1/4_%2B_1/8_%2B_1/16_%2B_⋯ https://en.wikipedia.org/wiki/1/2_%2B_1/4_%2B_1/8_%2B_1/16_%...
- juleiie 2mo agoThe point is that people who want to be secure can be more secure than ever while people who don’t care about security will be less secure than now. Author claims something slightly different but that’s how I would look at it. The extremes are more potent because either you get super secure 10 times vetted system or you get underpowered model slop with all the vulnerabilities it entails.
- Forgeties79 2mo agoThey’re not “bugs” they’re “quirks”! Our software is so quirky. It’s a feature!
- hollerith 2mo agoQuirky and adorable!
- thinkthatover 2mo agoDisagree, and in a way it feels like we are dealing with inverse issues: the security "skill" is well defined and will be also engaged with by an agent. Communication companies are further incentivized as any failure is at best reputational harm. Meanwhile SaaS companies are not strictly required to have good UX for their human end users, largely because those users will likely work around the issue. also network effect vs low costs of switching for for comms
- tptacek 2mo agoOne way to resolve the tension here is to note that CNE and lawful-intercept access to phones depends generally on platform vulnerabilities, not application code vulnerabilities. Low-level platform code churns less, absorbs more fixes under AI workloads than it does new features, and works in a constrained space where guardrails are easier to provide (and where those guardrails already have institutional support at Apple and Google). Over the long term this state of play could change, and IC/LEO organizations could start leaning more on application vulnerabilities than on platform RCEs. But the action would probably still coalesce around a couple of app-layer targets that could themselves be hardened.
- schoen 2mo agoI was hoping that the basebands and firmwares would get formally verified. Maybe they will ... with AI-written proofs!
- hashstring 2mo agoEh, we’ve seen many footholds of documented chains start in application code throughout time. To list a few: - WebKit (CVE-2016-4657) - WhatsApp (CVE-2019-3568) - Chrome (CVE-2021-38003) Then, also the distinction between platform vs application gets fuzzy with bugs in PassKit etc.
- tptacek 2mo agoWebKit and Chrome are platform codebases. WhatsApp is an example of what I'm talking about, though.
- leonidasrup 2mo agoLawful-intercept is in general implemented at the server equipment of telecom operator, not in the phones. In general, intelligence services don't access phones, if the interception can be done at the operator, because this could alert the subject or expose their techniques. According to a friend, who is working for an telecom company in Europe, Lawful-intercept is done at a server rack provided by and managed by a law enforcement agency, all traffic of this telecom company is copied to this server rack. The server rack also has a second direct connection to the law enforcement agency. Employees of the telecom company don't know which phone numbers are targeted or which traffic is intercepted.
- marcus_holmes 2mo ago> I see companies adding more sloppily written features with AI I think this is a side-effect of the old product management process adapting to AI. We (as an industry) were never very good at defining features rigorously, because there was a smart human in the loop who had to implement the feature and could push back on sloppy definitions. Whereas security bugs are easy for the LLM to define and fix.
- amarant 2mo agoMy hot take is that AI is a multiplier. Software engineers skill can be measured on a scale from -10 to +10, where 0 means you introduce as many bugs as you solve, or something along those lines(this scale is loosely defined, don't think too much about it) Any engineer who's skill value evaluates below 0 on my scale, ends up with a large negative number when they use AI. Anyone with a positive number ends up with a large positive number. The extra bugs you're seeing are from devs on the wrong side of 0 on my poorly defined scale.
- doginasuit 2mo ago> My guess is that we're getting better at finding _existing_ security issues with AI (and thus fixing those issues), but simultaneously adding more insecure surface areas _at a faster rate_. This does seem to be true in the vibe coding era, which I expect will implode at some point. But LLMs could certainly lead to a future where vulnerabilities are scarce. The best time to vet security factors is designing the model and writing the initial code, and LLMs are extraordinarily useful for this too. Most code with security implications is not written by a security expert. An LLM can be the most anal and well informed security expert you can find. If you write the code yourself but have one in the loop from the start, the code will be in much better shape.