13 ms·
I used o3 to find a remote zeroday in the Linux SMB implementation
- zison 1y agoVery interesting. Is the bug it found exploitable in practice? Could this have been found by syzkaller?
- mdaniel 1y agoI case anyone else didn't recognize that word: https://github.com/google/syzkaller https://github.com/google/syzkaller
- zielmicha 1y ago(To be clear, I'm not the author of the post, the title just starts with "How I")
- mdaniel 1y agoNoteable: > o3 finds the kerberos authentication vulnerability in 8 of the 100 runs And I'd guess this only became a blog post because the author already knew about the vuln and was just curious to see if the intern could spot it too, given a curated subset of the codebase
- moyix 1y agoHe did do exactly what you say – except right after that, while reviewing the outputs, he found that it had also discovered a different 0day.
- PunchyHamster 1y agoNow the question is whether spending same time to analyze that bit of code instead of throwing automated intern at it would be time spent better
- lyu07282 1y agoThe time they didn't spend reading the 13k LOCs themselves would've been time spent better. What?
- deleted 1y ago[deleted]
- Retr0id 1y agoThe article cites a signal to noise ratio of ~1:50. The author is clearly deeply familiar with this codebase and is thus well-positioned to triage the signal from the noise. Automating this part will be where the real wins are, so I'll be watching this closely.
- deleted 1y ago[deleted]
- tough 1y agoI was thinking about this the other day, wouldn't it be feasible to make fine-tune or something like that into every git change, mailist, etc, the linux kernel has ever hard? Wouldn't such an LLM be the closer -synth- version of a person who has worked on a codebase for years, learnt all its quirks etc. There's so much you can fit on a high context, some codebases are already 200k Tokens just for the code as is, so idk
- sodality2 1y agoI'd be willing to bet the sum of all code submitted via patches, ideas discussed via lists, etc doesn't come close to the true amount of knowledge collected by the average kernel developer's tinkering, experimenting, etc that never leaves their computer. I also wonder if that would lead to overfitting: the same bugs being perpetuated because they were in the training data.
- andix 1y ago1:50 is a great detection ratio for finding a needle in a haystack.
- epolanski 1y agoI don't think the author agrees as he points out the bugs weren't that difficult to find.
- Aachen 1y agoNah. I'm not an expert code auditor myself but I've seen my colleagues do it and I've seen ChatGPT try its hand. Even when I give it a specific piece of code and probe/hint in the right direction, it produces five paragraphs of vulnerabilities, none of which are real, while overlooking the one real concern we identified You can spend all day reading slop or you can get good at this yourself and be much more efficient at this task. Especially if you're the developer and know where to look and how things work already, catching up on security issues relevant to your situation will be much faster than looking for this needle in the haystack that is LLM output
- Hilift 1y agoDoes the vulnerability exist in other implementations of SMB?
- p_ing 1y agoImplementations of SMB (Windows, Samba, macOS, ksmbd) are going to be different (macOS has a terrible implementation, even though AFP is being deprecated). At this level, it's doubtful that the code is shared among all implementations.
- logifail 1y agoMy understanding is that ksmbd is a kernel-space SMB server "developed as a lightweight, high-performance alternative" to the traditional (user-space) Samba server... Q1: Who is using ksmbd in production? Q2: Why?
- pixl97 1y agoI would assume for the reason of being lightweight and high performance?
- foobar10000 1y agoSmb over 25gbit networks - user space samba is much worse there.
- Henchman21 1y agoThis is interesting to me! I regularly deploy 25G network connections, but I don’t think we’d run SMB over that. I am super curious the industry and use case if you’re willing to share!
- hackernudes 1y ago"SMB Direct" is RDMA based and ksmbd supports it. Samba does not. Disclaimer: I have not used it but was looking it up just yesterday.
- Henchman21 1y agoAppreciated, thank you.
- tinco 1y agoI ran SMB over a 20gbit network (2x 10gbit). The use case was 3D rendering (photogrammetry specifically). There were multiple render nodes, and a central service coordinating the rendering process. The projects would be on SSD's on the central SMBD server, and after they were manually configured (using Agisoft Metashape) they'd be rendered. Projects would sometimes start as tens of gigabytes worth of photos, and the artifacts (including intermediates) would balloon into the hundreds of gigabytes, we'd have dozens of these projects per week. I researched quite extensively prior to landing on SMB, but it really seems like there isn't a better way of doing this. The environment was mixed windows/linux, but if there was a better pure linux solution I would've pushed our office staff to switch to Ubuntu.
- iandanforth 1y agoThe most interesting and significant bit of this article for me was that the author ran this search for vulnerabilities 100 times for each of the models. That's significantly more computation than I've historically been willing to expend on most of the problems that I try with large language models, but maybe I should let the models go brrrrr!
- roncesvalles 1y agoA lot of money is all you need~
- b112 1y agoA lot of burned coal, is what. The "don't blame the victim" trope is valid in many contexts. This one application might be "hackers are attacking vital infrastructure, so we need to fund vulnerabilities first". And hackers use AI now, likely hacked into and for free, to discover vulnerabilities. So we must use AI! Therefore, the hackers are contributing to global warming. We, dear reader, are innocent.
- Balooga 1y agoBetween $3k and $30k to solve a single ARC-AGI problem [1]. Not sure if "100 runs" makes this comparable. [1] https://techcrunch.com/2025/04/02/openais-o3-model-might-be-costlier-to-run-than-originally-estimated/ https://techcrunch.com/2025/04/02/openais-o3-model-might-be-...
- mcbuilder 1y agoI think it gave up trying to solve Pokemon. :) Seriously, aren't these ARC-AGI problems easy for most people? They usually involve some sort of pattern recognition and visual reasoning.
- sdoering 1y agoSo basically running a microwave for about 800 seconds, or a bit more than 13 minutes per model? Oh my god - the world is gonna end. Too bad, we panicked because of exaggerated energy consumption numbers for using an LLM when doing individual work. Yes - when a lot of people do a lot of prompting, these 0ne tenth of a second to 8 seconds of running the microwave per prompt adds up. But I strongly suggest, that we could all drop our energy consumption significantly using other means, instead of blaming the blog post's author about his energy consumption. The "lot of burned coal" is probably not that much in this blog post's case given that 1 kWh is about 0.12 kg coal equivalent (and yes, I know that we need to burn more than that for 1kWh. Still not that much, compared to quite a few other human activities. If you want to read up on it, James O'Donnell and Casey Crownhart try to pull together a detailed account of AI energy usage for MIT Technology Review.[1] I found that quite enlightening. [1]: https://www.technologyreview.com/2025/05/20/1116327/ai-energy-usage-climate-footprint-big-tech/ https://www.technologyreview.com/2025/05/20/1116327/ai-energ...
- mezyt 1y agoMeanwhile, as a maintainer, I've been reviewing more than a dozen false positives slop CVEs in my library and not a single one found an actual issue. This article's is probably going to make my situation worse.
- SamuelAdams 1y agoMaybe, but the author is an experienced vulnerability analyst. Obviously if you get a lot of people who have no experience with this you may get a lot of sloppy, false reports. But this poster actually understands the AI output and is able to find real issues (in this case, use-after-free). From the article: > Before I get into the technical details, the main takeaway from this post is this: with o3 LLMs have made a leap forward in their ability to reason about code, and if you work in vulnerability research you should start paying close attention. If you’re an expert-level vulnerability researcher or exploit developer the machines aren’t about to replace you. In fact, it is quite the opposite: they are now at a stage where they can make you significantly more efficient and effective.
- tecleandor 1y agoNot even that. The author already knew the bug was there, and fed the LLM just the files related to the bug, with the explanation on how the methods worked and where to search, and even then, only 1 out of 100 times did it find the bug.
- sweetjuly 1y agoThere are two bugs in the article: one the author previously knew about and was trying to rediscover as an exploration as well as a second the author did not know about and stumbled into. The second bug is novel, and is what makes the blog post interesting.
- baq 1y agoprobably not. o3 is not free to use.
- jobswithgptcom 1y agoWow, interesting. I been hacking a tool called https://diffwithgpt.com https://diffwithgpt.com with a similar angle but indexing git changelogs with qwen to have it raise risks for backward compat issues, risks including security when upgrading k8s etc.
- empath75 1y agoGiven the value of finding zero days, pretty much every intelligence agency in the world is going to be pouring money into this if it can reliably find them with just a few hundred api calls. Especially if you can fine tune a model with lots of examples, which I don't think open ai, etc are going to do with any public api.
- treebeard901 1y agoYeah, the amount of engineering they have around controlling (censoring) the output, along with the terms of service, creates an incentive to still look for any possible bugs, but not allow it in the output. Certainly for Govt agencies and others this will not be a factor. It is just for everyone else. This will cause people to use other models and agents without these restrictions. It is safe to assume that a large number of vulnerabilities exist in important software all over the place. Now they can be found. This is going to set off arms race game theory applied to computer security and hacking. Probably sooner than expected...
- akomtu 1y agoThis made me think that the near future will be LLMs trained specifically on Linux or another large project. The source code is a small part of the dataset fed to LLMs. The more interesting is runtime data flow, similar to what we observe in a debugger. Looking at the codebase alone is like trying to understand a waterfall by looking at equations that describe the water flow.
- baq 1y agoit needs to be trained on on enough TLA+ traces, too.
- KTibow 1y ago> With o3 you get something that feels like a human-written bug report, condensed to just present the findings, whereas with Sonnet 3.7 you get something like a stream of thought, or a work log. This is likely because the author didn't give Claude a scratchpad or space to think, essentially forcing it to mix its thoughts with its report. I'd be interested to see if using the official thinking mechanism gives it enough space to get differing results.
- gizmodo59 1y agoHaving tried both I’d say o3 is in a league of it’s own compared to 3.7 or even Gemini 2.5 pro. The benchmarks may show not a lot of gain but that matters a lot when the task is very complex. What’s surprising is that they announced it last November and only now it’s released a month back now? (I’m guessing lots of safety took time but no idea). Can’t wait for o4!
- dieortin 1y agoAll your content threads from the past months consist on you saying how much better OpenAI products are than the competition, so that doesn’t inspire a ton of trust.
- gizmodo59 1y agoBecause in my use cases they are? Coding and math, science research are my primary use cases and codex with o3 and o3 consistently outperforms others in complex tasks for me. I can’t say a model is better just to appeal to HN. If another model is as good as o3 id use that in a second.
- sothatsit 1y agoI also feel similarly. o3 feels quite distinct in what it is good at compared to other models. For example, I think 2.5 Pro and Claude 4 are probably better at programming. But, for debugging, or not-super-well-defined reasoning tasks, or even just as a better search, o3 is in a league of its own. It feels like it can do a wider breadth of tasks than other models.
- nxobject 1y agoA small thing, but I found the author's project-organization practices useful – creating individual .prompt files for system prompt, background information, and auxiliary instructions [1], and then running it through `llm`. It reveals how good LLM use, like any other engineering tool, requires good engineering thinking – methodical, and oriented around thoughtful specifications that balance design constraints – for best results. [1] https://github.com/SeanHeelan/o3_finds_cve-2025-37899 https://github.com/SeanHeelan/o3_finds_cve-2025-37899
- kweingar 1y agoHow do we benchmark these different methodologies? It all seems like vibes-based incantations. "You are an expert at finding vulnerabilities." "Please report only real vulnerabilities, not any false positives." Organizing things with made-up HTML tags because the models seem to like that for some reason. Where does engineering come into it?
- deleted 1y ago[deleted]
- nindalf 1y agoThe author is up front about the limitations of their prompt. They say > In fact my entire system prompt is speculative in that I haven’t ran a sufficient number of evaluations to determine if it helps or hinders, so consider it equivalent to me saying a prayer, rather than anything resembling science or engineering. Once I have ran those evaluations I’ll let you know.
- 0points 1y agoAuthor seems to downplay their own expertise and attribute it to the LLM, while at the same time admitting he's vibe prompting the LLM and dismissing wrong results while hyping the ones that happen to work out for him. This seems more like wishful thinking and fringe stuff than CS.
- 1y ago
- dehrmann 1y agoAre there better tools for finding this? It feels like the sort of thing static analysis should reliably find, but it's in the Linux kernel, so you'd think either coding standards or tooling around these sorts of C bugs would be mature.
- grg0 1y agoNot the expert in the area, but "classic static analysis" (for lack of a better term) and concurrency bugs doesn't really check. There are specific modeling tools for concurrency, and they are an entirely different beast than static analysis that requires notation and language support to describe what threads access what data when. Concurrency bugs in static analysis probably requires a level of context and understanding that an LLM can easily churn through.
- yellow_lead 1y agoSome static analysis tools can detect use after free or memory leaks. But since this one requires reasoning about multiple threads, I think it would've been unlikely to be found by static analysis.
- firesteelrain 1y agoI really hope this is legit and not what keeps happening to curl [1] https://daniel.haxx.se/blog/2024/01/02/the-i-in-llm-stands-for-intelligence/ https://daniel.haxx.se/blog/2024/01/02/the-i-in-llm-stands-f...
- deleted 1y ago[deleted]
- ape4 1y agoSeems we need something like kernel modules but with memory protection
- martinald 1y agoI think this is the biggest alignment problem with LLMs in the short term imo. It is getting scarily good at this. I recently found a pretty serious security vulnerability in an open source very niche server I sometimes use. This took virtually no effort using LLMs. I'm worried that there is a huge long tail of software out there which wasn't worth finding vulnerabilities in for nefarious means manually but if it was automated could lead to really serious problems.
- tekacs 1y agoThe (obvious) flipside of this coin is that it allows us to run this adversarially against our own codebases, catching bugs that could otherwise have been found by a researcher, but that we can instead patch proactively.\ I wouldn't (personally) call it an alignment issue, as such.
- tekacs 1y agoA few days later, case in point (I'm in no way affiliated): https://news.ycombinator.com/item?id=44117465 https://news.ycombinator.com/item?id=44117465
- Legend2440 1y agoIf attackers can automatically scan code for vulnerabilities, so can defenders. You could make it part of your commit approval process or scan every build or something.
- martinald 1y agoA lot of this code isn't updated though. Think of how many abandoned wordpress plugins there are (for example). So the defenders could, but how do they get that code to fix it? I agree after time you end up with a steady state but in the short medium term the attackers have a huge advantage.
- bongodongobob 1y agoIt's a moot point unless attackers have better LLMs don't have access to.
- dboreham 1y agoI feel like our jobs are reasonably secure for a while because the LLM didn't immediately say "SMB implemented in the kernel, are you f-ing joking!?"
- simonw 1y agoThere's a beautiful little snippet here that perfectly captures how most of my prompt development sessions go: > I tried to strongly guide it to not report false positives, and to favour not reporting any bugs over reporting false positives. I have no idea if this helps, but I’d like it to help, so here we are. In fact my entire system prompt is speculative in that I haven’t ran a sufficient number of evaluations to determine if it helps or hinders, so consider it equivalent to me saying a prayer, rather than anything resembling science or engineering. Once I have ran those evaluations I’ll let you know.
- davidgerard 1y agoThis is just fuzzing with extra power consumption?
- brokensegue 1y agoCan you reconstruct finding this bug with traditional fuzzing?
- fsckboy 1y ago>It is interesting by virtue of being part of the remote attack surface of the Linux kernel. ...if your linux kernel has ksmbd built into it; that's a much smaller interest group
- theptip 1y agoThis is a great case study. I wonder how hard o3 would find it to build a minimal repro for these vulns? This would of course make it easier to identify true positives and discard false positives. This is I suppose an area where the engineer can apply their expertise to build a validation rig that the LLM may be able to utilize.
- eqvinox 1y agoAnyone else feel like this is a best case application for LLMs? You could in theory automate the entire process, treat the LLM as a very advanced fuzzer. Run it against your target in one or more VMs. If the VM crashes or otherwise exhibits anomalous behavior, you've found something. (Most exploits like this will crash the machine initially, before you refine them.) On one hand: great application for LLMs. On the other hand: conversely implies that demonstrating this doesn't mean that much.
- paulddraper 1y agohttps://security.googleblog.com/2024/11/leveling-up-fuzzing-finding-more.html?m=1 https://security.googleblog.com/2024/11/leveling-up-fuzzing-...
- ngneer 1y agohttps://news.ycombinator.com/item?id=42017771 https://news.ycombinator.com/item?id=42017771 Meh.
- paulddraper 1y agoThat seems to be really preoccupied with who was first, without looking at the magnitude of the results, which is far from "meh."
- ngneer 1y agoI think it was more a PoC. I would be more impressed if it was deployed in production. "we want to reiterate that these are highly experimental results". If the dividends are massive, would they not deploy it in production and tell the world about it?
- eqvinox 1y agoI mean, yes, they're doing it, but my question was really whether people share my belief that it's a particularly well-fitting application ;) (Also yeah feels like the "FIRST!!1!eleven" thing metastasized from comment sections into C-level executives…)
- baby 1y agoI have a presentation here on doing it to target zk bugs https://youtu.be/MN2LJ5XBQS0?si=x3nX1iQy7iex0K66 https://youtu.be/MN2LJ5XBQS0?si=x3nX1iQy7iex0K66
- mptest 1y agohttps://youtu.be/MN2LJ5XBQS0 https://youtu.be/MN2LJ5XBQS0 rest of the link is tracking to my (limited) understanding
- baby 1y agoPosted that in a haste, but meant to share as this might be interesting to people who are trying to do the same kind of things :) I have more updates now, reach out if you wanna talk!
- gerdesj 1y agoI'll have to get my facts straight but I'm pretty sure that ksmbd is ... not used much (by me). https://lwn.net/Articles/871866/ https://lwn.net/Articles/871866/ This is also nothing to do with Samba which is a well trodden path. So why not attack a codebase that is rather more heavily used and older? Why not go for vi?
- usr1106 1y agoGood link. After reading this it's not a surprise that this code has security vulnerabilities. But of course from knowing that there must be more to actually finding it, it's still a big leap. 4 years after the article, does any relevant distro have that implementation enabled?
- meander_water 1y agoI'm not sure about the assertion that this is the first vulnerability found with an LLM. For e.g. OSS-Fuzz [0] has found a few using fuzzing, and Big Sleep using an agent approach [1]. [0] https://security.googleblog.com/2024/11/leveling-up-fuzzing-finding-more.html?m=1 https://security.googleblog.com/2024/11/leveling-up-fuzzing-... [1] https://googleprojectzero.blogspot.com/2024/10/from-naptime-to-big-sleep.html?m=1 https://googleprojectzero.blogspot.com/2024/10/from-naptime-...
- seanheelan 1y agoIt's certainly not the first vulnerability found with an LLM =) Perhaps I should have been more clear though. What the post says is "Understanding the vulnerability requires reasoning about concurrent connections to the server, and how they may share various objects in specific circumstances. o3 was able to comprehend this and spot a location where a particular object that is not referenced counted is freed while still being accessible by another thread. As far as I'm aware, this is the first public discussion of a vulnerability of that nature being found by a LLM." The point I was trying to make is that, as far as I'm aware, this is the first public documentation of an LLM figuring out that sort of bug (non-trivial amount of code, bug results from concurrent access to shared resources). To me at least, this is an interesting marker of LLM progress.
- fHr 1y agomeanwhile boomers out here still thinking they are better than AI wehen even local gemma3 models can write better code then them allready
- tomalbrc 1y ago“I brute forced an ai to help me find potential zero day bugs”
- mettamage 1y agoI wonder how often it will say there’s a vulnerability where there is non. Running it 100 times is a lot
- stonepresto 1y agoI know there were at least a few kernel devs who "validated" this bug, but did anyone actually build a PoC and test it? It's such a critical piece of the process yet a proof of concept is completely omitted? If you don't have a PoC, you don't know what sort of hiccups would come along the way and therefore can't determine exploitability or impact. At least the author avoided calling it an RCE without validation. But what if there's a missing piece of the puzzle that the author and devs missed or assumed o3 covered, but in fact was out of o3's context, that would invalidate this vulnerability? I'm not saying there is, nor am I going to take the time to do the author's work for them, rather I am saying this report is not fully validated which feels like a dangerous precedent to set with what will likely be an influential blog post in the LLM VR space moving forward. IMO the idea of PoC || GTFO should be applied more strictly than ever before to any vulnerability report generated by a model. The underlying perspective that o3 is much better than previous or other current models still remains, and the methodology is still interesting. I understand the desire and need to get people to focus on something by wording it a specific way, it's the clickbait problem. But dammit, do better. Build a PoC and validate your claims, don't be lazy. If you're going to write a blog post that might influence how vulnerability researchers conduct their research, you should promote validation and not theoretical assumption. The alternative is the proliferation of ignorance through false-but-seemingly-true reporting, versus deepening the community's understanding of a system through vetted and provable reports.
- lyu07282 1y agoAre you saying you want PoCs that trigger a crash from the use-after-free or you would only be satisfied by full on RCE PoCs?
- stonepresto 1y agoPoCs should at least trigger a crash, overwrite a register, or have some other provable effect, the point being to determine: 1) If it is actually a UAF or if there is some other mechanism missing from the context that prevents UAF. 2) The category and severity of the vulnerability. Is it even a DoS, RCE, or is the only impact causing a thread to segfault? This is all part of the standard vulnerability research process. I'm honestly surprised it got merged in without a PoC, although with high profile projects even the suggestion of a vulnerability in code that can clearly be improved will probably end up getting merged.
- jp0001 1y agoWe followed a very similar approach at work, created a test harness and tested all the models available in AWS bedrock and the OpenAI. We created our own code challenges not available on the Internet for training with vulnerable and non-vulnerable inline snippets and more contextual multi-file bugs. We also used 100 tests per challenge - I wanted to do 1000 test per challenge but realized that these models are not even close to 2 Sigma in accuracy! Overall we found very similar results. But, we were also able to increase accuracy using additional methods - which comes as additional costs. The issue I see overall is that we found is when dealing with large codebases you'll need to put blinders on the LLMs to shorten context windows so that hallucinated results are less likely to happen. The worst thing would be to follow red herrings - perhaps in 5 years we'll have models used for more engineering specific tasks that can be rated with Six Sigma accuracy if posed with the same questions and problems sets.
- bandrami 1y agoThe blinders give you a problem in that a lot of security issues aren't at a single point in the code but at where two remote points in the code interact.
- jp0001 1y agoCorrect. Dynamic runtime interactions will always be a hard problem as it’s hard to see in static code even for humans.
- qoez 1y agoThis is why AI safety is going to be impossible. This easily could have been a bad actor who would use this finding for nefarious acts. A person can just lie and there really isn't any safety finetuning that would let it separate the two intents.
- resiros 1y agoI think an approach like AlphaEvolve is very likely to work well for this space. You've got all the elements for a successful optimization algorithm: 1) A fast and good enough sampling function + 2) a fairly good energy function. For 1) this post shows that LLMs (even unoptimized) are quite good at sampling candidate vulnerabilities in large code bases. A 1% accuracy rate isn't bad at all, and they can be made quite fast (at least very parallelizable). For 2) theoretically you can test any exploit easily and programmatically determine if it works. The main challenge is getting the energy function to provide gradient—some signal when you're close to finding a vulnerability/exploit. I expect we'll see such a system within the next 12 months (or maybe not, since it's the kind of system that many lettered agencies would be very interested in).
- geraneum 1y agoThis has become a common recurrence recently. Have a problem with clear definition and evaluation function. Let LLM reduce the size of solution space. LLMs are very good at pattern reconstruction, and if the solution has a similar pattern to what was known before, it can work very well. In this case the problem is a specific type of security vulnerability and the evaluator is the expert. This is similar in spirit to other recent endeavors where LLMs are used in genetic optimization; on a different scale. Here’s an interesting read on “Mathematical discoveries from program search with large language models” which was I believe was also featured in HN the past: https://www.nature.com/articles/s41586-023-06924-6 https://www.nature.com/articles/s41586-023-06924-6 One small note, concluding that the LLM is “reasoning” about code just _based on this experiment_ is bit of a stretch IMHO.
- 1oooqooq 1y agoi can ask offline o3 about that cve and get a reply, does that mean the author used a model that knew about the vulnerability?
- antirez 1y agoEither I'm very lucky or as I suspected Gemini 2.5 PRO can more easily identify the vulnerability. My success rate is so high that running the following prompt a few times is enough: https://gist.github.com/antirez/8b76cd9abf29f1902d46b2aed3cdc1bf https://gist.github.com/antirez/8b76cd9abf29f1902d46b2aed3cd...
- mehulashah 1y agoThe scary part of this is that the bad guys are doing the same thing. They’re looking for zero day exploits, and their ability to find them just got better. More importantly, it’s now almost automated. While the arms race will always continue, I wonder if this change of speed hurts the good guys more than the bad guys. There are many of these, and they take time to fix.
- jokoon 1y agoWow I think the NSA already has this, without the need for a LLM.
- curtisszmania 1y ago[dead]