12 ms·
Mythos Finds a Curl Vulnerability
- ahofmann 5mo agoPutting on my tinfoil-hat: Sooo, the guy who runs the test and delivers the report could just have removed the more interesting bugs and delivered those to any three letter agency?
- bilekas 5mo ago[flagged]
- Ekaros 5mo agoCurl is likely one of the very much more combed over pieces of code at this point. It feels like it has some special draw for people looking for vulnerabilities. Not that it doesn't mean some novel idea can't be looked or checked still.
- cakealert 5mo ago> No, based on cURL's history, it really seems like they would love to have found a really novel bug. You just confirmed that you didn't read the article. "Eventually, I was instead offered that someone else, who has access to the model, could run a scan and analysis on curl for me using Mythos and send me a report."
- bilekas 5mo agoI'm not sure how that proves I didn't read the article ?
- croon 5mo agoSomeone external to the curl team ran the test. If that third party found a severe CVE that they could use across all the global curl attack surface, and did not disclose it back to the curl team, the third party could keep using the exploit until discovered independently.
- AnssiH 5mo agoThe test was run by an unnamed third party, so cURL's history has no relevance to their benevolence.
- casey2 5mo agocurl's source is public so what would be the gain in the rigmarole? Now if the prompt was "create a patch that inserts a zero-day while fixing a bug" that would be impressive.
- rzmmm 5mo agoQuote: "My personal conclusion can however not end up with anything else than that the big hype around this model so far was primarily marketing. I see no evidence that this setup finds issues to any particular higher or more advanced degree than the other tools have done before Mythos. Maybe this model is a little bit better, but even if it is, it is not better to a degree that seems to make a significant dent in code analyzing." It's a good reminder for us all that the competition in this space is rough and lots of more or less subtle marketing is involved.
- greendude29 5mo agoI'd go out and say the marketing is not subtle. The hype and fanboys/girls are so in line with the marketing that any level of skepticism is seen a an act of defection, but if you look at the words, hyperbole and volume that is used, there is nothing subtle about it. It's almost Trump-esque - "this model will change everything forever; we are doomed; we are saved; we will all be fired; we will all be rich", etc
- xantronix 5mo agoThat's a pretty good encapsulation of the parallels between the political and the technological: One necessarily thrives upon the other and are inextricable. This moment is a culmination of all the disenfranchisement the bodypolitik have suffered, looking for any possible means of escape or elevation. AI and Trumpism, for their own respective cohorts, are salvation, on offer by different frontmen but ultimately in service of the same system. They need the hype to pay off way more than we do. So many of us who still write code directly stand to lose nothing of our capabilities if the marketing claims cannot hold water.
- ehnto 5mo agoI seem to be totally outside the hype bubble, but I have to suspect there is a lot of imagineering and wild extrapolations in the elss technical hype bubbles. I am curious but no enough to go looking.
- 5mo ago
- bilekas 5mo ago> The single confirmed vulnerability is going to end up a severity low CVE planned to get published in sync with our pending next curl release 8.21.0 in late June My mind still cannot understand the quality and refinement that's gone into cURL. It really is the perfect example of something done so right, that people barely think twice about.
- dotancohen 5mo agoCurl and SQLite are my favourite examples of properly engineered, rigourously tested _anything_. It's really philosophical - those projects' contribution requirements demand such rigor, and the maintainers stand by that demand. A non-load-bearing document (not project code) is what makes that possible - very reminiscent of Einstein's thought experiments leading to tangible projects such as GPS or Descartes's belief that all problems can be solved through rational thinking.
- ontouchstart 5mo agoSome people must be working on training some models exclusively on high quality OSS code base like curl and SQLite without the noise of low quality training data. I would do that with 100% local models from scratch.
- pjmlp 5mo agoEasy, it shows what is achievable if there is a high bar for quality in every single line of code that gets commited, reviewed and merged, regardless of the programming language. However in the days of race to bottom, offshoring for penies, and now LLM powered code generation, this is a quality most companies won't care unless there is liability in place.
- bilekas 5mo ago> Easy, it shows what is achievable if there is a high bar for quality in every single line of code that gets commited This is becoming a more and more overlooked/underrated feature. I genuinely believe it would be impossible in any company that depends on shareholder value. I am yet to convince any company I've worked in without bloody hands that we need to solve old tech debt and refactor certain things etc.
- yjftsjthsd-h 5mo ago> The source code consists of 660,000 words, which is 12% more words than the entire English edition of the novel War and Piece. Typo, or is there a spoof I should go read?
- dotancohen 5mo agoPerhaps he was dictating. Does it say anything else? Just 'Aaaarggghhhh'?
- Hamuko 5mo agoDoubt it considering that Daniel Stenberg is Swedish. English dictation when you speak English as a second language with an accent is quite annoying.
- Tistron 5mo agoVoice input works really well for people speaking English with a Swedish accent. I think the accent of most educated Swedes is mostly a case of prosody. For sure there are some sounds we say slightly differently than native English speakers. We often have some trouble with /s/ and /z/, but I don't know, "war and peace", I think that's easily understood. Source: voice typing this with Swedish vocal chords, and only had to correct "different lives" to "differently", and add /[^\w\s]/.
- aitchnyu 5mo agoAndroid voice input works with kids using both English and native words, here in India. The country runs schools in 25+ primary languages, each with dialects, so a TV/phone with voice input is more marvelous than the nitpicks discussed here.
- dotancohen 5mo agoI understand completely. You don't want to know what the machine produced, when I asked it for "a new display".
- deleted 5mo ago
- yjftsjthsd-h 5mo ago> Not particularly “dangerous” I'm not sure that follows. As noted, curl was already analyzed to death with every tool available; most software isn't at that level.
- bilekas 5mo agoI don't think I understand what you mean, the "not particularly dangerous" comment was in relation to the vulnerability that was found right ? Surely they would know what constitutes a lower severity level.
- vidarh 5mo agoThe "not particularly dangerous" is a headline for a section talking about Mythos, not the vulnerability.
- bilekas 5mo agoAh okay, that makes a bit more sense. I read it wrong. Then the comment is absolutely fair.
- Ekaros 5mo agoMy guess is that it is in category of "you are holding it wrong". Still worth fixing, but requires very specific user input for example. Or very weird scenario. Or in some less used protocol or flag combination.
- croon 5mo agoSure, but isn't it a verdict on Mythos compared to other models? If so, it would still follow. "Most software" isn't analyzed as much as curl, by either other tooling or other models, that might well find close to the same as Mythos did. As such, Mythos then isn't especially/particularly dangerous.
- anygivnthursday 5mo agoBut Mythos is not marketed as a tool that can do the same as other tools already available maybe slightly better, but as a revolution.
- AntiUSAbah 5mo agoThere is always marketing involved and people should be able to put marketing into perspective. Also curl in this regard is a open source project, relativly small but critical, well known and used everywhere. Besides image libraries, tools like curl or sudo, su, passwd, etc. would also be my first try. Mythos is still not known at all what it can do. What does it mean from cost and benchmark pov to have a 10 Trillion parameter model? Nonetheless, the fact that LLMs got significant better in finding this, better than humans, started to happen half a year ago? so at one point we need to address the elefant in the room and state that today you need to do security scanning additional with LLMs. You need to take this serious. In worst case, use Anthropics marketing to state that its a must now and something changed.
- flohofwoe 5mo ago> Nonetheless, the fact that LLMs got significant better in finding this, better than humans, started to happen half a year ago? *rolls eyes* regular static analyzers also have been "better than humans" for decades, being better than a human at a specific mechanical task really doesn't mean much. The interesting new thing is the type of potential "fuzzy bugs" described in the article that LLMs are able to identify (a comment not matching the code it describes, uncommon usage of a 3rd party library, mismatch of code and a protocol it implements, or often just generally weird looking code somebody should have a closer look at... this closes a gap in the traditional debugging toolboxes, but shouldn't replace them)
- AntiUSAbah 5mo agoYou don't have to dismantle a comment on a microlevel. It has been clear for ages that certain type of bugs or issues are better solved from software. But there was still plenty of things a proper SecOps Person would be able to find with help from tooling which automatic tooling wouldn't find. Taking a limited amount of resources and focusing on the critical things. I do think this is gone now. Same with Threat modeling etc.
- pixl97 5mo agoStatic analyzers are balls. For every real bug they find you are dealing with with piles of false positives and negatives. Now, I'm not saying you shouldn't use them. They do catch the low hanging fruit. It's that LLMs actually have a much better understanding of things like intent when looking at your code and general architecture configurations that can lead to problems. As you say we've had static analyzers forever, hence why they aren't dropping out 50 new CVE's a day. LLMs are. There is a massive stack of software out there that is getting analyzed and exploited at a rate faster than it's getting patched. Adding to that things like NPMs exploited package of the day and popular github repository takeovers this year looks massively different from last year in quantity and quality of exploits alone.
- mohsen1 5mo agoI don't know about Mythos but in recent weeks I've noticed Opus is constantly failing to fix things in tsz[0] vs GPT 5.5 can easily churn out fixes that are solid and pass tests. I've stopped paying for Claude for now and all my money is going to OpenAI at the moment. Either Opus is massively nerfed or GPT 5.5 is really head and shoulder higher in terms of very difficult tasks. The last percent of conformance tests in tsz are really really difficult and I've seen Opus bailing again and again. So annoying to waste time and tokens to finally get "this is too involved" or "this requires a multi-week sprint to fix". [0] https://tsz.dev https://tsz.dev
- dyauspitr 5mo agoHaving never used Claude and only Codex, does Claude actually say “this is too involved” as a response to a prompt?
- mohsen1 5mo agoYes it does. Usually after hours of working and not getting results
- redditor98654 5mo agoI am curious, what kind of work do you use Claude for that sometimes requires hours of working. In my case, I have never seen it go off for more than 10 mins and even that is very rare.
- mohsen1 5mo agohttps://github.com/mohsen1/tsz https://github.com/mohsen1/tsz
- big_youth 5mo agodebugging code. I had some issue so I create a plan to root cause that would run the code, change some functions or variables and run again until we get a confirmed answer. I just work up to that very workflow this morning. I ran last night and finished at around 3am with ~200k tokens spent. Fixed the issue and created a follow up doc for things that it could not verify.
- perching_aix 5mo agoIt's a shame he seems to reject the idea of actually diving in and using these tools interactively: > It’s not that I would have a lot of time to explore lots of different prompts and doing deep dive adventures anyway. His expertise I think would elevate the results quite a bit. Although if he never uses LLMs, which it reads like he doesn't, I guess it might backfire just as well. Prompting style (still?) does matter after all, certainly in my experience anyways.
- jph00 5mo agoHe states in the article that they use LLMs for this purpose and find them extremely useful.
- perching_aix 5mo agoWhich can be true without this also being true: > using these tools interactively I did read the article. It seems to me they're using LLMs in a prepared manner instead, as mere scanners that produce reports.
- SpicyLemonZest 5mo agoPerhaps I'm misreading something? From my reading of the article, it doesn't sound like Anthropic offered to let him use Mythos in any other way than that.
- perching_aix 5mo agoHe explains in the article that he failed to actually secure access in the end, even though it was approved. Someone else prompted the model on his behalf, and just passed on the findings.
- OtherShrezzing 5mo agoHe posts about his use of language models a lot on Mastodon[0]. He does lots with language models, but doesn't buy all the way into the hype. I'd say he's one of. most reasonable & balanced voices on the subject of AI use in software today. Happy to use the technology, more than willing to push back on marketing bs. [0] https://mastodon.social/@bagder https://mastodon.social/@bagder
- almogodel 5mo ago[flagged]
- absynth 5mo agoI routinely used to compile C programs on other compilers to find defects that one or another didn't find. Compiling on Windows vs Linux. You could summarize / minimize it down to compiling it with warning as errors etc but you'd be missing the point. The point wasn't actual cross-platform portability even though that was a nice side effect. It was to flush out all the weird edge cases. Edges like security flaws. Buffer overflows are usually platform specific. There are plenty of other ways to find these issues but simply recompiling for a different platform surfaces all sorts of issues.
- apexalpha 5mo ago> An amazingly successful marketing stunt for sure. This. Well done by Antropic. It even reached the CISO of my small semi-government org in the Netherlands, who slightly panicked at the announced 'tsunami' of vulnerabilities that was coming with Mythos. Got us some more money and priority with the board, though. Never waste a good marketing scare.
- fpesce 5mo agoI don't agree with the "no tsunami in sight": if you don't look at 100+ bugs in Firefox and many more OSS projects, bunch of old unseen-before OpenBSD/Linux RCEs, and a few LPE in just 2 or 3 weeks for Linux itself... IMO, this does not sound like marketing scare, there is spike of vulnerability disclosures - high quality, low false positives - that can be sensed... It feels like we're speedrunning through few-years worth of high quality bug reports in just a few weeks.
- apexalpha 5mo agoThe LPEs were not found with Mythos but with existing, publicly available models.
- stingraycharles 5mo agoAnd also: they did an earlier run with Opus to discover bugs (like segfaults). In February, Opus discovered a whole bunch of security related bugs, but didn’t exploit them. Mythos, in turn, was fed these bugs and told to exploit them. Not saying it’s not impressive, but it was literally told “here are all the places our metal detector says there may be gold, please find gold”.
- dralley 5mo agoThere is a significant difference between being able to see one flaw and being able to chain together multiple disparate flaws, to be fair.
- Aurornis 5mo ago
- utopiah 5mo agoWon my bet "voted 10 [vulnerabilities] but in retrospect as you are familiar with Claude and such tooling if you already used any of recent model to done some kind of security review then I'd drop to 1 or even 0." https://mastodon.pirateparty.be/@utopiah/116537456780283420 https://mastodon.pirateparty.be/@utopiah/116537456780283420
- jongjong 5mo agoI'm looking forward to trying Mythos run against my 5000-line, instant-finality, quantum-resistant blockchain project and decentralized exchange (an additional 5000 lines). I already ran all the models up to Opus 4.6 and they couldn't find anything.
- deleted 5mo ago[deleted]
- nevi-me 5mo ago> These tools and the analyses they have done have triggered somewhere between two and three hundred bugfixes merged in curl through-out the recent 8-10 months or so. If you've just gone through a lengthy analysis of your code with other AI tools, surely it's reasonable not to expect to see hundreds more from a new tool? It should be possible, unless more bugs are introduced, to eventually get to a state where there are no more bugs in your code. Process aside, it sounds like Daniel expected to find dozens/hundreds more bugs.
- jaapz 5mo agoMythos was kind of hyped as the tool that would discover much more bugs than any currently available tool
- pbmonster 5mo agocurl had ~15 CVEs in 2026 so far. You surely don't think those (and the one Mythos found) were the last security bugs still left in the code base? There certainly will be more, in fact Daniel predicts ~50 CVEs for the entire year. But Mythos found 1. After all that hype. 1.
- knowaveragejoe 5mo agoMaybe curl is just... better hardened? Firefox posted hundreds in April.
- pbmonster 5mo agoThat's not the argument. Yes, curl is insanely hardened. But still, they currently have a new CVE every couple of weeks. Mythos didn't accelerate this much, no more than all the other AI-assisted security analysis they've been doing anyway. Which either means that, tragically for Mythos, it only got to analyze the code base just after ALL the bugs where finally ironed out and now curl is bug free forever after - or Mythos isn't really all that good, dozens/hundreds more bugs remain and will be found in the next months and years. I just think the former is a bit unlikely.
- jedisct1 5mo agoSwival found many more vulnerabilities without Mythos https://github.com/swival/security-audits https://github.com/swival/security-audits
- NitpickLawyer 5mo agoWhat's going on in this thread? It's weird how prevalent the negativity towards mythos is, and I'm not sure if it's people throwing the baby out with the bathwater or something more tinfoil-adjacent coordinated campaign. I also noticed this on a thread a few days ago, before the mozilla post. There were dozens of comments saying basically "mythos is vaporware". I get the idea that they're using it for marketing. Of course they are. But to reduce it at "just marketing" feels either ill informed or outright wrong. Unless you have reasons to not believe the dozens of credentialed, well respected people in the field that have already shared their opinions after working with mythos. Plenty of them on all the social media sites. And then there's the team at mozilla. They wrote a blog about this, and they've worked with anthropic before, using opus 4.6 and found and fixed 22 vulnerabilities. Then they worked with mythos and found and fixed 271 vulnerabilities. Unless you're going to accuse them of being shills, these are unquestionable numbers. The model is quantitatively better at this thing. And it matches what everyone is saying. I think there are better things to accuse anthropic of, than that they are simply lying for marketing purposes. Of course they'll use this as a marketing campaign, but there's plenty of evidence out there that there is something there, that the model is simply better than previous generations at this. Don't fall for the cheap reductionist stuff, just because you don't like them, or feel that this is marketing fluff. It doesn't feel like a gimmick, even if it gets used to push their agenda. Something, something, propaganda often uses true statements as well.
- countWSS 5mo agoHere and on reddit, AI debugging is viewed as some weird shallow pattern-matching that obviously fails to spot real stuff and overload the maintainers. Instead of getting to "spotless record" of zero flaws, the people start rationalizing that "X is not a real bug" and inventing justifications for their(obviously bad) code, which is critique they can't accept from AI, only through human debate that they can't close with a WONTFIX. Once the bug is actually usable, the tune changes completely.
- mschuster91 5mo ago> Here and on reddit, AI debugging is viewed as some weird shallow pattern-matching that obviously fails to spot real stuff and overload the maintainers. That's because that is what a lot of people did in the last years [1] to pad their resumes or to force developers to backport patches to older (but supported) kernel versions that wouldn't have gone in if they didn't have a CVE attached [2]. Maintainers have been legitimately swamped with low-quality spam for a very long time. Only recently, in the last few months, AI actually got "good enough", the problem is that maintainers still have to differentiate between AI slop by wannabes and by AI-assisted reports reviewed and refined by actual human professionals. [1] https://www.zdnet.com/article/how-fake-security-reports-are-swamping-open-source-projects-thanks-to-ai/ https://www.zdnet.com/article/how-fake-security-reports-are-... [2] https://opensourcewatch.beehiiv.com/p/linux-gets-cve-security-business https://opensourcewatch.beehiiv.com/p/linux-gets-cve-securit...
- andromaton 5mo agoIf priced like other Anthropic models, Mythos will make vulnerability discovery a lot more accessible. The author compares it to AISLE, ZeroPath, and OpenAI’s Codex Security. AISLE and ZeroPath are much more expensive. OpenAI’s Codex Security is gated. Most people don't care about the first two and don't complain about the latter's policy because they are all specialized models and/or harnesses. Mythos will be available to all.
- vibedev999 5mo ago> AISLE and ZeroPath are much more expensiv AISLE is *cheaper* for sure
- Semkas 5mo agoI'm disinclined to be overly generous to Antrophic, but I have to say that regardless of whether the talk of Mythos being uniquely dangerous was mostly cynical: It would be great if this starts a trend of giving security-critical software a few months head start with any new significantly improved model.
- sdhrag 5mo ago[dead]
- jrflo 5mo agoI know that the Mythos hype is part marketing by anthropic, but isn't it possible that with a highly scrutinized codebase, there just aren't any notable security exploits in it's current state? The fact that it found nothing isn't necessarily an incrimination against it, especially when other tools had identified hundreds of exploits previously. Seems like it's been completely picked over (for now).
- billyoneal 5mo agoPeople lost their minds over the mythos announcement specifically because they found something in FreeBSD, which had a reputation as being one of those picked over code bases.
- srcreigh 5mo agoI can't help but think that curl is, by nature, a relatively simple and well-contained tool. Compare to an operating system or web browser or database or billion dollar company codebase. It makes some sense that Mythos/ChatGPT 5.5 might be that much better with complexities that curl just doesn't have because it's a basic tool. Like yeah curl is obviously extremely fully featured as an "anything client" but it's orders of magnitude less complex than other software we rely on.
- joelthelion 5mo agoI agree it's rather basic but as stated in the article, its code is still longer than war and peace. There is still plenty of opportunities for security vulnerabilities in something of that size.
- sausagefeet 5mo agoCurl is a lot more complicated than, I believe, you think. Most people know of it simply as a CLI to hit an HTTP(S) endpoint and write it out. But: 1. It supports basically any file transfer protocol. 2. It is a library that is designed for long running processes. 3. Because it's designed for long running processes, it makes use of every trick it can to pipeline and re-use connections and resources. 4. It has an asynchronous API so it can be integrated into any existing event loop. Is a web browser or database more complicated? Most certainly, they solve really massive problems. But curl is certainly more complicated than probably most application code that uses it.
- breakpointalpha 5mo agoFrom the post: "curl is currently 176,000 lines of C code when we exclude blank lines. The source code consists of 660,000 words, which is 12% more words than the entire English edition of the novel War and Peace. ... curl is installed in over twenty billion instances. It runs on over 110 operating systems and 28 CPU architectures. It runs in every smart phone, tablet, car, TV, game console and server on earth." I wouldn't call that simple or well contained... Most OS or web browsers don't run on cars or tvs.
- 5mo ago
- romaniv 5mo ago"I signed the contract for getting access, but then nothing happened. Weeks went past and I was told there was a hiccup somewhere and access was delayed. Eventually, I was instead offered that someone else, who has access to the model, could run a scan and analysis on curl for me using Mythos and send me a report. To me, the distinction isn’t that important." Really? We're talking about (essentially) a product demo from a trillion dollar industry fueled by debt. Clearly, blog posts like this have an immense influence on the perception of usefulness of the particular model and AI in general. With so much staked on this for the company, wouldn't you want to be sure that you're using the actual product without anyone messing with the results in any way?
- nottorp 5mo ago> (I am purposely leaving out the identity of the individual(s) involved in getting the curl analysis done as it is not the point of this blog post.) I would very much like to know if they were independent or affiliated to Anthropic. > My personal conclusion can however not end up with anything else than that the big hype around this model so far was primarily marketing. ... because of this.
- tgtweak 5mo agoI feel like, if it was a codebase without using any security analysis tools, there would have been some more significant findings - perhaps they can re-run it on an 18 month old commit and see how many it found that were subsequenty found and fixed? Anyway, I think the case that frontier and next-gen models will get increasingly adept at finding vulnerabilities and that those on the receiving end of those vulnerabilities need to be on top of it.
- ostif-derek 5mo agoUnfortunately that doesn't help much. LLMs are really really good at digging up known vulns, so much so that they often falsely declare known vulns as new and novel ones. They have the CVEs in their training data, know how to look up ossfuzz logs, etc.
- readthenotes1 5mo agoKinda burying the lede: AI tools found over a dozen CVEs in curl last year, and hundreds of bugs. "Primarily AISLE, Zeropath and OpenAI’s Codex Security have been used to scrutinize the code with AI. These tools and the analyses they have done have triggered somewhere between two and three hundred bugfixes merged in curl through-out the recent 8-10 months or so. A bunch of the findings these AI tools reported were confirmed vulnerabilities and have been published as CVEs. Probably a dozen or more."
- toraway 5mo agoNot exactly "burying the lede" since Daniel already posted an update about it months ago [1] with extensive discussion in numerous of articles [2] including on this site [3]. [1] https://lists.haxx.se/pipermail/daniel/2025-September/000127.html https://lists.haxx.se/pipermail/daniel/2025-September/000127... [2] https://www.theregister.com/software/2025/10/02/curl-project-swamped-with-ai-slop-finds-not-all-ai-is-bad/588432 https://www.theregister.com/software/2025/10/02/curl-project... [3] https://news.ycombinator.com/item?id=45449348 https://news.ycombinator.com/item?id=45449348
- brunoborges 5mo agoAI not finding a security issue on cURL has more to do with lack of widespread security issues than the model's capacity of finding them.
- patrickmeenan 5mo agoAs far as I can tell, the messaging around Mythos is that it takes the expertise of the top security experts and top-level language, protocol and code experts and makes that available to anyone with access. The danger was in giving that access to the world before the defenders had access to that level of expertise. Curl HAS had security, protocol and language experts poking at it for years because of how central it is to everything. That Mythos found anything is interesting but not a sign that it's been marketing hype and isn't dangerous. You can bet that 99.99% of projects aren't nearly as secure as curl and it doesn't matter if they are open or closed source (LLM's will happily decompile closed-source projects and explore). Unless your project has been fuzzed and gone over with existing AI tooling and by experts, expect that it can already be hacked - even with the tooling that is out there now and that something like Mythos makes it accessible for an even wider population pool with less expertise to use.
- 2001zhaozhao 5mo agoTake my upvote. Anthropic never claimed superhuman performance, only speed and scale. That it doesn't find much in terms of new vulnerabilities in a well-studied piece of software says nothing about its overall potential for dangerous misuse.
- EMM_386 5mo agoIf an AI agent finds zero bugs in a software utility, how can that be viewed in the sense the AI agent is not very good at finding bugs? What if there are actually zero bugs? > Five issues felt like nothing as we had expected an extensive list. The expectation here may not match reality, but not necessarily because Mythos isn't as capable as claimed. curl may just happen to be a well-hardened tool that doesn't have too many security vulnerabilities in its present state.
- zamadatix 5mo agoThe author considered the same w.r.t. remaining bugs: > More to find > These were absolutely not the last bugs to find or report. Just while I was writing the drafts for this blog post we have received more reports from security researchers about suspected problems. The AI tools will improve further and the researchers can find new and different ways to prompt the existing AIs to make them find more. > We have not reached the end of this yet. > I hope we can keep getting more curl scans done with Mythos and other AIs, over and over until they truly stop finding new problems. And that makes sense, it'd be quite the argument of coincidence to say there was just 1 proper find remaining & it was only Mythos that managed to find it just at the point in time it released while the other projects have been hoovering up every other find quickly until that point. Possible, but not the safest assumption to start questioning with.
- AtNightWeCode 5mo agoShould have scanned it with Mythos on an older code base before all these other sec issues was resolved with other tools. Or use the other tools to introduce the same kind of errors in other parts of the code base to see if Mythos would have found it. A problem is that these tools seems smarter than they are cause they already read seen the answer key.
- deleted 5mo ago[deleted]
- jorisw 5mo agoLove Daniel's writing style here. Fact based, concise, easy to read
- plexescor 5mo agoI personally belive its a marketing stunt and they are just using actual humans to find the bugs/vulns
- tedd4u 5mo agoIt's also a convenient way to get press (and investor valuation) for a new model with releasing it (word is they don't have enough hardware to do so).
- deferredgrant 5mo ago[flagged]
- vb-8448 5mo agoI guess we miss fundamental information: how much in terms of time and token usage took to the "middle guy" to create the report? Next question: could it be that OP can use Mythos in a better way since he knows better the project?
- deleted 5mo ago[deleted]
- hmokiguess 5mo ago> curl is one of the most fuzzed and audited C codebases in existence (OSS-Fuzz, Coverity, CodeQL, multiple paid audits). Finding anything in the hot paths (HTTP/1, TLS, URL parsing core) is unlikely. The way this reads sounds more like the LLM dismissed trying rather than it tried and failed, I've seen Claude do that often unless I probe it to challenge itself, curious here what actually happened.
- theaniketmaurya 5mo agoWho is using Mythos to find these things and where do they run it?
- ilia-a 5mo agoIMHO Mythos was more of a marketing ploy. When it comes to security and AI, all top tier publicly accessible models (GPT 5.5, Opus 4.7) and even near-top like Deepseek 4 PRO can do a very good job given detailed harness on how to spot issues and cross-validate them to avoid false positives.
- tuananh 5mo agoInteresting. curl team found that Mythos is mostly hype while Calif team found Mythos amazing. I would think Calif (a security firm) is a better team to better utilize such tool.
- 23aqsI 5mo agoLike clockwork, criticism of the Alpha Omega apparatchiks is flagged. They know how to protect their income streams while open source authors get nothing.
- selectedambient 5mo agosooo tldr; curl is safe.