13 ms·
Potential issues in curl found using AI assisted tools
https://joshua.hu/llm-engineer-review-sast-security-ai-tools-pentesters https://joshua.hu/llm-engineer-review-sast-security-ai-tools...
https://joshua.hu/files/AI_SAST_PRESENTATION.pdf https://joshua.hu/files/AI_SAST_PRESENTATION.pdf
- deleted 1y ago[deleted]
- simonw 1y agoHere are 55 closed PRs in the curl repo which credit "sarif data" - I think those are the ones Daniel is talking about here https://github.com/curl/curl/pulls?q=is%3Apr+sarif+is%3Aclosed https://github.com/curl/curl/pulls?q=is%3Apr+sarif+is%3Aclos... This is notable given Daniel Stenberg's reports of being bombarded by total slop AI-generated false security issues in the past: https://www.linkedin.com/posts/danielstenberg_hackerone-curl-activity-7324820893862363136-glb1 https://www.linkedin.com/posts/danielstenberg_hackerone-curl... Concerning HackerOne: "We now ban every reporter INSTANTLY who submits reports we deem AI slop. A threshold has been reached. We are effectively being DDoSed. If we could, we would charge them for this waste of our time" Also this from January 2024: https://daniel.haxx.se/blog/2024/01/02/the-i-in-llm-stands-for-intelligence/ https://daniel.haxx.se/blog/2024/01/02/the-i-in-llm-stands-f...
- octocop 1y agoThe models used have improved quite well since then, I guess his change of opinion shows that.
- simonw 1y agoI think it's more about how people are using it. An amateur who spams him with GPT-5-Codex produced bug reports is still a waste of his time. Here a professional ran the tools and then applied their own judgement before sending the results to the curl maintainers.
- tptacek 1y agoI keep irritating people with this observation but this was the status quo ante before AI, and at least an AI slop report shows clear intent; you can ban those submitters without even a glance at anything else they send.
- sumeno 1y agoThe current scale of poor reports was absolutely not the status quo before AI
- tptacek 1y agoThe last time I was staffed on a project that had to do this, we were looking at many dozens per day, virtually all of them bogus, many attached to grifters hoping to jawbone the triage person into paying a nominal fee to get them to shut up. It would be weird if new tooling like LLMs didn't accelerate it, but that's all I'd expect it to do.
- whizzter 1y agoIt's probably also the difference of idiots hoping to cash out/get credit for vulnerabilities by just throwing ChatGPT at the wall compared to this where it seems a somewhat seasoned researcher is trialing more customized tools.
- Twirrim 1y agoNo, he's still dealing with a flood of crap, even in the last few weeks, off more modern models. It's primarily from people just throwing source code at an LLM, asking it to find a vulnerability, and reporting it as-read, without having any actual understanding of if it is or isn't a vulnerability. The difference in this particular case is it's someone who is: 1) Using tools specifically designed for security audits and investigations. 2) Takes the time to read and understand the vulnerability reported, and verifies that it is actually a vulnerability before reporting. Point 2 is the most significant bar that people are woefully failing to meet and wasting a terrific amount of his time. The one that got shared from a couple of weeks ago https://hackerone.com/reports/3340109 https://hackerone.com/reports/3340109 didn't even call curl. It was straight up hallucination.
- tomjakubowski 1y agoSome of those bugs, like using the wrong printf-specifier for a size_t, would be flagged by the compiler with the right warning flags set. An AI oracle which tells me, "your project is missing these important bug-catching compiler warning flags," would be quite useful. A few of these PRs are dependabot PRs which match on "sarif", I am guessing because the string shows up somewhere in the project's dependency list. "Joshua sarif data" returns a more specific set of closed PRs. https://github.com/curl/curl/pulls?q=is%3Apr+Joshua+sarif+data+is%3Aclosed+ https://github.com/curl/curl/pulls?q=is%3Apr+Joshua+sarif+da...
- flohofwoe 1y agoThis is exactly what I'd want from an 'AI coding companion'. Don't write or fix the code for me (thanks but I can manage that on my own with much less hassle), but instead tell me which places in the code look suspicious and where I need to have a closer look. When I ask Claude to find bugs in my 20kloc C library it more or less just splits the file(s) into smaller chunks and greps for specific code patterns and in the end just gives me a list of my own FIXME comments (lol), which tbh is quite underwhelming - a simple bash script could do that too. ChatGPT is even less useful since it basically just spend a lot of time to tell me 'everything looking great yay good job high-five!'. So far, traditional static code analysis has been much more helpful in finding actual bugs, but static analysis being clean doesn't mean there are no logic bugs, and this is exactly where LLMs should be able to shine. If getting more useful potential-bugs-information from LLMs requires an extensively customized setup then the whole idea is getting much less useful - it's a similar situation to how static code analysis isn't used if it requires extensive setup or manual build-system integration instead of just being a button or menu item in the IDE or enabled by default for each build.
- walthamstow 1y agoCursor BugBot is pretty good for this, we did the free trial and it was so popular with our devs that we ended up keeping it. Occasional false positives aside, it's very useful. It saves time for both the PR submitter and the reviewer.
- simonw 1y agoSuggestion: run a regex to remove those FIXME comments first, then try the experiment again. I often use Claude/GPT-5/etc to analyze existing repositories while deliberately omitting the tests and documentation folders because I don't want them to influence the answers I'm getting about the code - because if I'm asking a question it's likely the documentation has failed to answer it already!
- trenchpilgrim 1y agoI use Zed's "Ask" mode for this all the time. It's a read only mode where the LLM focuses on figuring out the codebase instead of modifying it. You can toggle it freely mid conversation.
- chmod775 1y agoI really didn't expect a story about curl and AI to be positive for once. Some history: https://hn.algolia.com/?q=curl+AI https://hn.algolia.com/?q=curl+AI
- dwedge 1y agoYeah this is really fair play to Daniel Stenberg that he still approached these AI generated bug reports with an open mind after all the problems he's had.
- athorax 1y agoI think the big difference is that these aren't AI generated bug reports. They are bugs found with the assistance of AI tools that were then properly vetted and reported in a responsible way by a real person.
- Legion 1y agoBasically using AI the way we have used linters and other static analysis tools, rather than thinking it's magic and blindly accepting its output.
- stocksinsmocks 1y agoIn the defense of the language models, the bugs were written by humans in the first place. Human vetting is not much of a defense.
- NegativeK 1y ago> Human vetting is not much of a defense. The issue I keep seeing with curl and other projects is that people are using AI tools to generate bug reports and submitting them without understanding (that's the vetting) the report. Because it's so easy to do this and it takes time to filter out bug report slop from analyzed and verified reports, it's pissing people off. There's a significant asymmetry involved. Until all AI used to generate security reports on other peoples' projects is able to do it with vanishingly small wasted time, it's pretty assholeish to do it without vetting.
- 1970-01-01 1y agoNotice it was 'a set of tools' They're using it correctly. It's a system of tools, not an autopilot.
- ragnese 1y agoWell, that's how Mr. Stenberg described it, but he wasn't the one using them. I don't know how the contributor feels about his AI tool(s).
- mos_basik 1y agoI haven't read it yet, but later in the mastodon thread, stenberg says "this is [the contributor's] (long) blog post on his work: https://joshua.hu/llm-engineer-review-sast-security-ai-tools-pentesters https://joshua.hu/llm-engineer-review-sast-security-ai-tools...".
- tptacek 1y agoIt's weird that the discussion has collapsed down to "autopilots" vs. "abstention". I'm thrilled to be converging on an understanding that it instead "people who understand what they're trying to do" vs. "vibe coders".
- nxobject 1y agoIn defense of the cynics, I get the impression in a situation where (a) there's a lot of company marketing hype in such a competitive market that begs cynicism, and (b) we're constantly learning the boundary of trained LLMs can actually do (and can't), as well as unusual emergent workflows, that really do make a difference.
- Timshel 1y agoI did not read it, but this article from the contributor should contain more details: https://joshua.hu/llm-engineer-review-sast-security-ai-tools-pentesters https://joshua.hu/llm-engineer-review-sast-security-ai-tools... (mentioned in https://mastodon.social/@bagder/115241413210606972 https://mastodon.social/@bagder/115241413210606972).
- bgwalter 1y agoIf something is found by Valgrind, we can reproduce it ourselves. Here we get private bug reports found by "his set of AI assisted tools". The set seems to be: https://joshua.hu/llm-engineer-review-sast-security-ai-tools-pentesters https://joshua.hu/llm-engineer-review-sast-security-ai-tools... So he likes ZeroPath. Does that get us any further? No, the regular subscription costs $200 and the free one-time version looks extremely limited and requires yet another login. Also of course, all low hanging fruit that these tools detect will be found quickly in open source (provided that someone can afford a subscription), similar to the fact that oss-fuzz has diminishing returns.
- simonw 1y agoPresumably the bug reports were private because some of them might relate to curl security. You can see the fixes that resulted from this in the PRs that mention "sarif" in the curl repository: https://github.com/curl/curl/pulls?q=is%3Apr+sarif+is%3Aclosed https://github.com/curl/curl/pulls?q=is%3Apr+sarif+is%3Aclos...
- cubefox 1y agoThis should probably link to the original blog post by Joshua Rogers: https://joshua.hu/llm-engineer-review-sast-security-ai-tools-pentesters https://joshua.hu/llm-engineer-review-sast-security-ai-tools... ("Hacking with AI SASTs: An overview of 'AI Security Engineers' / 'LLM Security Scanners' for Penetration Testers and Security Teams")
- tempodox 1y agoNow that is how LLM assistance for coding can be useful. Would be interesting to know which set of tools was used exactly. How might one reproduce this kind of assistance for other code bases?
- simonw 1y agoSee Joshua's post for details: https://joshua.hu/llm-engineer-review-sast-security-ai-tools-pentesters https://joshua.hu/llm-engineer-review-sast-security-ai-tools... Tools included ZeroPath, Corgea and Almanax.
- excitedrustle 1y ago[dead]
- konart 1y agoLove this one: https://mastodon.social/@icing@chaos.social/115244064143435733 https://mastodon.social/@icing@chaos.social/1152440641434357... >tldr >The code was correct, the naming was wrong.
- alganet 1y agoSomething sounds fishy in this. Has these bugs really been found by AI? (I don't think they were). If you read Corgea's (one of the products used) "whitepaper", it seems that AI is not the main show: > BLAST addresses this problem by using its AI engine to filter out irrelevant findings based on the context of the application. It seems that AI is being used to post-process the findings of traditional analyzers. It reduces the amount of false positives, increasing the yield quality of the more traditional analyzers that were actually used in the scan. Zeropath seems to use similar wording like "AI-Enabled Triage" and expressions like "combining Large Language Models with AST analysis". It also highlights that it achieves less false positives. I would expect someone who developed this kind of thing to setup a feedback loop in which the AI output is somehow used to improve the static analysis tool (writing new rules, tweaking existing ones, ...). It seems like the logical next step. This might be going on on these products as well (lots of in-house rule extensions for more traditional static analysis tools, written or discovered with help of AI, hence the "build with AI" headline in some of them). Don't get me wrong, this is cool. Getting an AI to triage a verbose static analysis report makes sense. However, it does not mean that AI found the bugs. In this model, the capabilities of finding relevant stuff are still capped at the static analyzer tools. I wonder if we need to pay for it. I mean, now that I know it is possible (at least in my head), it seems tempting to get open source tools, set them to max verbosity, and find which prompts they are using on (likely vanilla) coding models to get them to triage the stuff.
- simonw 1y agoLooks like you're reacting to the Hacker News title here, which is currently " Daniel Stenberg on 22 curl bugs found by AI and fixed" That's an editorialized headline (so it may get fixed by dang and co) - if you click through to what Daniel Stenberg said he was more clear: > Joshua Rogers sent us a massive list of potential issues in #curl that he found using his set of AI assisted tools. AI-assisted tools seems right to me here.
- alganet 1y agoIf the title changes, it is still a valid critique of the tools, how they might work, and a possible way of getting them for free. Also, think about it: of course I read Joshua's report. Otherwise, how could I have known the names of the products he used?
- silisili 1y ago> I have already landed 22(!) bugfixes thanks to this, and I have over twice that amount of issues left to go through Sounds like it was a lot more than 22, assuming most are valid.
- retr0reg 1y agoI work in a ML security R&D startup called Pwno, we been working on specifically putting LLMs into memory security for the past year, we've spoken at Black Hat, and we worked with GGML (llama.cpp) on providing a continuous memory security solution by multi-agents LLMs. Somethings we learnt alone the way, is that when it comes to specifically this field of security what we called low-level security (memory security etc.), validation and debugging had became more important than vulnerability discovery itself because of hallucinations. From our trial-and-errors (trying validator architecture, security research methodology e.g., reverse taint propagation), it seems like the only way out of this problem is through designing a LLM-native interactive environment for LLMs, validate their findings of themselves through interactions of the environment or the component. The reason why web security oriented companies like XBOW are doing very well, is because how easy it is to validate. I seen XBOW's LLM trace at Black Hat this year, all the tools they used and pretty much need is curl. For web security, abstraction of backend is limited to a certain level that you send a request, it whether works or you easily know why it didn't (XSS, SQLi, IDOR). But for low-level security (memory security), the entropy of dealing with UAF, OOBs is at another level. There are certain things that you just can't tell by looking at the source but need you to look at a particular program state (heap allocation (which depends on glibc version), stack structure, register states...), and this ReACT'ing process with debuggers to construct a PoC/Exploit is what been a pain-in-the-ass. (LLMs and tool callings are specifically bad at these strategic stateful task, see Deepmind's Tree-of thoughts paper discussing this issue) The way I've seen Google Project Zero & Deepmind's Big Sleep mitigating this is through GDB scripts, but that's limited to a certain complexity of program state. When I was working on our integration with GGML, spending around two weeks on context, tool engineering can already lead us to very impressive findings (OOBs); but that problem of hallucination scales more and more with how many "runs" of our agentic framework; because we're monitoring on llama.cpp's main branch commits, every commits will trigger a internal multi-agent run on our end and each usually takes around 1 hours and hundreds of agent recursions. Sometime at the end of the day we would have 30 really really convincing and in-depth reports on OOBs, UAFs. But because how costly to just validate one (from understanding to debugging, PoC writing...) and hallucinations, (and it is really expensive for each run) we had to stop the project for a bit and focus solving the agentic validation problem first. I think when the environment gets more and more complex, interactions with the environment, and learning from these interactions will matters more and more.
- qustrolabe 1y agoOh so AI usage news could be positive after all. Not to undermine huge issue of slop reports spam, but I'm so happy to see something besides doomerism
- imiric 1y ago"AI" tools can be very powerful once you approach them as what they are: very good pattern matchers and generators. This ability far surpasses anything a human could do. Detecting potential issues in software is a great application of the technology. The key word is "potential", though. They're still wildly unpredictable and unreliable, which is why an expert human is required to validate their output. The big problem is the people overhyping the technology, selling it as "AI", and the millions deluded by the marketing. Amidst the false advertising, uncertainty, and confusion, people are forced to speculate about the positive and negative impacts, with wild claims at both extremes. As usual, the reality is somewhere in the middle.
- lowbloodsugar 1y agoWhen I read “we consider nread == 0 as reading a byte and we shouldn’t” I immediately think of all the things that look like bugs but are there because some critical piece of infrastructure relies on that behavior. AI isn’t going to know about that unless you tell it, and the problem is that there’s plenty of folks who have job security precisely because they don’t write that down.
- deleted 1y ago[deleted]
- dude250711 1y ago[flagged]
- kissgyorgy 1y ago[flagged]
- noelwelsh 1y agoMore interesting to me is how to stop these bugs from occurring in the first place. The example given in the thread is the kind of bug that C (and mutation) excels at creating.
- viraptor 1y agoThe linked blog post https://joshua.hu/llm-engineer-review-sast-security-ai-tools-pentesters https://joshua.hu/llm-engineer-review-sast-security-ai-tools... shows that most of the used tools can be run in ci and comment on the PRs.
- swaits 1y agoAnd how many would’ve been avoided by finishing the rust port?
- daxfohl 1y agoSo whereas previously, repo owners were getting flooded with AI-generated PRs that were complete slop, now they're going to be flooded with PRs that contain actual bugfixes. IDK which problem is worse!
- blixt 1y agoIt wasn't immediately obvious to me what the AI tools were? He mentioned that multiple other tools failed to find anything, so I'm very curious to hear what made this strategy so superior.
- NooneAtAll3 1y agothere's a blog link https://joshua.hu/llm-engineer-review-sast-security-ai-tools-pentesters https://joshua.hu/llm-engineer-review-sast-security-ai-tools... that has Products chapter I guess mastodon link is simply a confirmation that bugs were indeed bugs, even with wrong code snippets?
- srcreigh 1y agoLink should be updated to this https://joshua.hu/llm-engineer-review-sast-security-ai-tools-pentesters https://joshua.hu/llm-engineer-review-sast-security-ai-tools...
- dang 1y agoI've added that link to the toptext, but I can't quite tell which URL should be the starting point.
- tiahura 1y agoPerhaps Anthropic, OpenAI, and Google could compete by auditing and monitoring the top projects?
- runningmike 1y agoThere are some good SAST scanners and many bad commercial scanners. Many people advocate for the use of AI technology for SAST testing. There are even people and companies that deliver SAST scanners based on AI technology. However: Most are just far from good enough. In the best case scenario, you’ll only be disappointed. But the risk of a false sense of security is enormous. Some strong arguments against AI scanners can be found on https://nocomplexity.com/ai-sast-scanners/ https://nocomplexity.com/ai-sast-scanners/
- renox 1y agoOnce Claude found a bug in my code but I had to explain the structure of the data. Then and only then it found the bug.
- Eggpants 1y agoI have to admit, I expected a couple of "You should rewrite it in Rust" hipster posts by now... Maybe they caught on that those types of posts were not having the effect they thought they would? I kid, I kid... mostly
- redbell 1y agoSomehow related: You did this with an AI and you do not understand what you're doing here: https://news.ycombinator.com/item?id=45330378 https://news.ycombinator.com/item?id=45330378
- anonymars 1y agoYeah, I'm quite confused. See also https://news.ycombinator.com/item?id=44411185 https://news.ycombinator.com/item?id=44411185 ("AI slop security reports submitted to curl"), and especially https://news.ycombinator.com/item?id=43907376 https://news.ycombinator.com/item?id=43907376 ("Curl: We still have not seen a valid security report done with AI help")
- cadamsdotcom 1y agoAI is non-deterministic as we know. That makes its results unpredictable. So don’t have AI create your bugs. Instead have your AI look for problems - then have it create deterministic tools and let tools catch the issues in a repeatable, understandable, auditable way. Have it build short, easy to understand scripts you can commit to your repo, with files and line numbers and zero/nonzero exit codes. It’s that key step of transforming AI insights into detection tools that transforms your outcomes from probabilistic to deterministic. Ask it to optimize the tools so they run in seconds. You can leave them in the codebase forever as linters, integrate them in your CI, and never have that same bug again.
- mmsc 1y agoAlways fun to wake up (ok; I didn't wake up, I got off a 10 hour flight) to see my work on the front page of hn. I'll be doing a retrospective in a few weeks when the dust has settled, as well as new tools I've been made aware of.
- polycaster 1y agoI’m curious what tools that may be.
- gavinray 1y agoI thoroughly enjoyed the post, one of the few lengthier blog posts I read start-to-finish. Seems like ZeroPath might be worth looking into if the price is reasonable
- mmsc 1y agoThank you, it means a lot.
- yapyap 1y agoYes the AI gave him leads and a talented programmer still has to follow up on them one by one. It’s like a police facial recognition, they can help police but there is no way they are “replacing police”
- yapyap 1y agoAlso, what an intense presentation style. Red borders around every slide and very flashy images
- riedel 1y agoI am a bit worried about the abuse of those tools. I wonder if the tools have policies and mechanisms around those (no clue how, but like forced disclosure if they detect that scanned code is OSS, free usage for OSS teams). Seems otherwise like great tools to accelerate finding 0-zero days. Maybe ever worse, when building supply chain attacks, one could relatively easily test against the detection mechanism before contributing malicious code (wonder if they could detect and block malicious use/probing). However, i guess long term it will make our software more secure. I guess it is always an arms race.
- Michael_Keller 1y ago[dead]
- shivasurya 1y agoLove this take actually and have been working on this and published this way back 2023/2024. Recently, I've been inspired by Claude-code & Cline agentic flow + tool looping, I experimented the same with tools like file_read, dir_list and throwing in few sast tools, security prompts on Wordpress plugin ecosystem (say with 10k-100k active installation) and scanned around ~600 and to my surprise it yielded ~45 critical, ~120 high severity issues and accounting 20% for non-reachability vuln. Spent around 6$ and ~40 million tokens with grok-4 fast reasoning model and the results were impressive, I gave a try with claude-sonnet but significantly rate-limited despite having 50$ credits from anthropic for research. You can read about my experience here: https://codepathfinder.dev/blog/introducing-secureflow-cli-to-hunt-vuln/ https://codepathfinder.dev/blog/introducing-secureflow-cli-t... Old post: https://shivasurya.me/security-reviews/sast/2024/06/27/automate-security-code-reviews-with-cody-ai.html https://shivasurya.me/security-reviews/sast/2024/06/27/autom...
- appleaday1 1y agoTHe Canadian gubernment should probably get a bug bounty program so I can present some of my findings to the them that I found using ai and tested or mapped things out on some of their public facing apps on the app store/play store.
- deleted 1y ago[deleted]
- ronsor 1y agoTo anyone who thinks about our current situation for than a few minutes, AI is * Clearly useful to people who are already competent developers and security researchers * Utterly useless to people who have no clue what they're doing But the latter group's incompetency does not make AI useless in the same way that a fighter jet is not useless because a toddler cannot pilot it.
- Terr_ 1y agoThough it might be bad news for the companies that got big by saying they'd be able to make infinite money selling fighter-jets to all the children of the world.
- ranger_danger 1y ago> “A good tool in the hands of a competent person is a powerful combination,” says Daniel Stenberg.
- 0points 1y agoDaniel Stenberg has been vocal about AI generated patches in the past, and it's interesting to see him changing course here. https://news.ycombinator.com/item?id=38845878 https://news.ycombinator.com/item?id=38845878 https://news.ycombinator.com/item?id=43907376 https://news.ycombinator.com/item?id=43907376 https://media.ccc.de/v/froscon2025-3407-ai_slop_attacks_on_the_curl_project https://media.ccc.de/v/froscon2025-3407-ai_slop_attacks_on_t...
- ronsor 1y agoHe's been against people submitting garbage they don't understand, not AI as a whole.
- the_jeremy 1y ago> * Utterly useless to people who have no clue what they're doing I disagree. I'm making a board game of 6 colors of hexes, and I wanted to be able to easily edit the board. The first time around, I used a screenshot of a bunch of hexagons and used paint to color them (tedious, ugly, not transparent, poor quality). This time, I asked ChatGPT to make an SVG of the board and then make a JS script so that clicking on a hex could cycle through the colors. Easier, way higher quality, adjustable size, clean, transparent. It would've taken me hours to learn and set that up for myself, but ChatGPT did it in 10min with some back and forth. I've made one SVG in my life before this, and never written any DOM-based JS scripts. Yes, it's a toy example, but you don't have to knwo what you're doing to get useful things from AI.
- ambicapter 1y agoI'd love to know how an 'AI-native' SAST uses AI under the hood but I supposed that's a trade secret.
- deleted 1y ago[deleted]
- pwnfunction 1y agodo you think the "bug" exist in the latent space? i don't think all bugs does.. the bugs exist as long as variant exist in the trained weights. until we have some kinda rl env for verifying bugs.. its never gonna work "well".