5 ms·
It seems like bug hunting might be the one area where AI is actually making the world a better place.
by charonn0 3mo ago
It seems like bug hunting might be the one area where AI is actually making the world a better place.
- ashleyn 3mo agoHow many were introduced by misuse of AI coding/vibe coding though?
- DANmode 3mo agoHow many were known, and put on the roadmap because war got hot?
- iJohnDoe 3mo ago[flagged]
- stackghost 3mo agoAt Microslop? Evidently, lots.
- miffy900 3mo agohighly unlikely for many of them. SharePoint, bitlocker, Active directory, hyper-v, rdp, DHCP and MSMQ are all software/technologies that have decades of history and long pre-dated LLMs. seriously, do people not realise it was entirely possible to write insecure or bad code before LLMs?
- shakna 3mo agoSure, that's true. It is also true that Copilot is currently in use developing Bitlocker and Sharepoint. So I wouldn't be confident saying it was one or the other.
- pjmlp 3mo agoEspecially if they made heavy use of offshoring, which I would bet they did.
- warshinder 3mo agoIt’s like people don’t remember the whole outsourcing trend and all the awful code that came from that.
- datakan 3mo agoDon't forget the damned interns!
- mr_mitm 3mo agoEveryone but me!
- christophilus 3mo agoHey, I wrote slop at Microsoft way before it was cool.
- oskarw85 3mo ago[dead]
- matwood 3mo ago> seriously, do people not realise it was entirely possible to write insecure or bad code before LLMs? Some of these threads make me think every line of code written pre-LLMs was apparently perfect in all ways. Feels like romanticizing the past.
- dpoloncsak 3mo agoIt's funny because SharePoint and AD are so god-forsakenly awful you would think they're vibe coded if you didn't know any better
- ColdStream 3mo agoIt is hard to tell, the code may genuinely be decent quality or not. That is the issue with vibe coding. Increased output but reduced understanding. So if something does go wrong, one has to hope that there is still enough understanding to address it quickly.
- Leherenn 3mo agoFor what it's worth, at my workplace AI has uncovered quite a few issues that have been there for a decade or two and survived countless rounds of careful reviews, external security analysis, pen testing and so forth for all those years.
- damian260 3mo agoThe distinction I'd draw is between AI-assisted and AI-generated. Using AI to write isolated functions you understand and review is different from prompting your way to a complete system you can't debug. The second case is where you get surprising failures at runtime that no amount of linting catches.
- high_na_euv 3mo ago3 or 4
- onion2k 3mo agoVibe-coded apps probably have loads, but mostly because they're using less capable models than the people who're doing the bug-hunting. Once vibe-coders are using models like Mythos too you should expect the number of bugs in vibe-coded apps to collapse quickly, because the LLM will write the bugs but will also fix them (assuming the system prompt tells it to.)
- Foobar8568 3mo agoFor some reasons, models are blind to their own output...
- coldtea 3mo agoOne reason seems obvious/intuitive: because their own output matches their own biases, that is, is a direct result of their model walks.
- tyre 3mo agoI think this could be true if you use a single Claude context, since it has its own reasoning in the context. But a separate code review agent does much better, in my experience.
- tyre 3mo agoNot my experience with my homie Claudius. The code review agent usually has feedback to be resolved before committing, which includes bugs and unhandled edge cases. Sometimes the primary context is understandably embarrassed. Sometimes it truly be your own people.
- subscribed 3mo agoNot my experience. I use Claude to build me a small web app I needed for ages, I'm using its superpowers to discuss/talk about architecture /brainstorm/build plans, and always a fresh context with the built plans. A few times the plan implementer (subagent writing a code assigned to the task) found something lacking (more often it was test issue related to the code, not the actual flaw of the implementation plan), fixed it on the fly, or found something in the review (last step of each task). Also there's always a technical review for each "slice" (set of tasks), consisting of code review and e2e tests. Only when it passes I do the fresh code review of the changes with the Opus or Fable again. Happens rarely rarely, but it did found a few issues. Code works every time, I have yet to find the fault myself. Of course there are issues and I need to read the output especially in the planning phase very carefully and yes, Opus disappointed me many times trying to weasel out from something it "agreed with me" (and entered into architecture + todos)... Of course right after agreeing to use devcontainers it proceeded to attempt installing a handful of node modules in my os, so intended up running it in the bwrap (pain in the ass in itself). But it works, it's fascinating, and I have the app I actually needed. Not magical but useful.
- Razengan 3mo agoIs it worse than what humans were doing on their own anyway?
- prmoustache 3mo agoWell yes if they expect to find and correct more bugs every update which is pretty much what they are saying.
- Razengan 3mo agoIf AI ever manages to make something half as atrocious as Windows XP that would probably be the first proof of AGI.
- wccrawford 3mo agoNo, just more prolific.
- beebmam 3mo ago99.9% of people complaining about AI making the world a worse place would be fully happy with AI if they shared in the economic benefits of automation.
- azinman2 3mo agoNot if it ends up deskilling society and taking away what brings us meaning in life.
- deleted 3mo ago[deleted]
- ivell 3mo agoI don't see why AI would deskill what you love to do. People still do embroidery even though mass manufacturing exists. If you love something you would continue doing it irrespective of automation.
- ares623 3mo agoYes. I love eating pieces of string. I'm partial to wool myself, really filling. I suppose you're a cake enjoyer, miss Marie Antoinette?
- wood_spirit 3mo agoFor a lot of people it’s not the job that is rewarding it the role having a job gives them in life and home life. Financially contributing to the household through earning it through work is a meaningful and rewarding thing that can define the near total of how good you feel about yourself thing even if you don’t like the job?
- deleted 3mo ago[deleted]
- voidUpdate 3mo agoBut do people make a living doing embroidery by hand? Or is it more of a hobby?
- sublimefire 3mo agoIt was like a bunch of mythos scans through and through which then generated the reports for everyone to implement, not sure if in all orgs though. Mythos was great as it came from the top, i.e. a clear incentive. I think bug reporting otherwise does not reach the engineers unless an incident is raised by the customer care against a responsible team.
- amelius 3mo agoBad guys can use AI too. Not sure about the net result.
- subscribed 3mo agoIf your ally is stealing trade secrets from your companies and spying on your politicians and journalists, is it really a good guy? Sure there are _worse_ guys, but the supposed good ones aren't.
- ptdorf 3mo agoDepends. If the AI fixes don't introduce new bugs or unnecessary complexity (where bugs like to hide), then yes.
- kmfrk 3mo agoThere's also the weird scenario of a reporting bias where instead of admitting to a bunch of vulnerabilities, you can frame it as "look at how useful AI is". We even have companies implementing solutions for ffmpeg vulnerabilities themselves instead of just handing them the vulns to fix themselves. It's very possible the tide reverses if it's not in people's interest to advocate for AI anymore, so we better not get too used to it just in case. It's nice that we currently have an alignment of AI advocacy and infosec, though. Maybe Microsoft can even point their AI to their questionable UX and UI practices next.
- prmoustache 3mo agoI am curious, they say we should expect a trend of more security patches every update. It this is true and Microsoft devs are using agents it means that AI is doing a shitty job and is introducing more bugs/vulns per month than it fixes them. Otherwise you would have had a lot of bugs/vulns fixed in the first 2 security updates then a steady and significant reduction every month because any new code would have been scanned and fixed before release.