38 ms·
> Wasn't the scaffolding for the Mythos run basically a line of bash that loops through every file of the codebase and prompts the model to find vulnerabilities
by johnfn 6mo ago
> Wasn't the scaffolding for the Mythos run basically a line of bash that loops through every file of the codebase and prompts the model to find vulnerabilities in it? That sounds pretty close to "any gold there?" to me, only automated.
But the entire value is that it can be automated. If you try to automate a small model to look for vulnerabilities over 10,000 files, it's going to say there are 9,500 vulns. Or none. Both are worthless without human intervention.
I definitely breathed a sigh of relief when I read it was $20,000 to find these vulnerabilities with Mythos. But I also don't think it's hype. $20,000 is, optimistically, a tenth the price of a security researcher, and that shift does change the calculus of how we should think about security vulnerabilities.
- amazingamazing 6mo agoCitation needed for basically all of this. You basically are creating a double standard for small models vs mythos…
- johnfn 6mo agoThe citation is the Anthropic writeup.
- amazingamazing 6mo agoThey did not say what you are saying… > If you try to automate a small model to look for vulnerabilities over 10,000 files, it's going to say there are 9,500 vulns.
- johnfn 6mo agoWhat I am saying is that the approach the Anthropic writeup took and the approach Aisle took are very different. The Aisle approach is vastly easier on the LLM. I don't think I need a citation for that. You can just read both writeups. The "9500" quote is my conjecture of what might happen if they fix their approach, but the burden of proof is definitely not on me to actually fix their writeup and spend a bunch of money to run a new eval! They are the ones making a claim on shaky ground, not me.
- cycomanic 6mo agoSo you can't imagine anything between bruteforce scan the whole codebase and cut everything up in small chunks and scan only those? You don't think that security companies (and likely these guys as well) develop systems for doing this stuff? I'm not a security researcher and I can imagine a harness that first scans the codebase and describes the API, then another agent determines which functions should be looked at more closely based on that description, before handing those functions to another small llm with the appropriate context. Then you can even use another agent to evaluate the result to see if there are false positives. I would wager that such a system would yield better results for a much lower price. Instead we are talking about this marketing exercise "oohh our model is so dangerous it can't be released, and btw the results can't be independently verified either"
- user34283 6mo ago[dead]
- johnfn 6mo agoI explained why this won't work elsewhere in the thread[1]. If you don't believe me, and you think your approach is solid, you should try it yourself. It's only a couple of dollars, and it would be extremely popular -- just look at how popular this article, using improper methodology, was! Hey, maybe you're right, and you can prove us all wrong. But I'd bet you on great odds that you're not. [1]: https://news.ycombinator.com/item?id=47734710 https://news.ycombinator.com/item?id=47734710
- integralid 6mo ago>Or none We already know this is not true, because small models found the same vulnerability.
- deleted 6mo ago[deleted]
- tptacek 6mo agoNo, they didn't. They distinguished it, when presented with it. Wildly different problem.
- enraged_camel 6mo agoYeah. And it is totally depressing that this article got voted to the top of the front page. It means people aren’t capable of this most basic reasoning so they jumped on the “aha! so the mythos announcement was just marketing!!”
- BoiledCabbage 6mo ago> because small models found the same vulnerability. With a ton of extra support. Note this key passage: >We isolated the vulnerable svc_rpc_gss_validate function, provided architectural context (that it handles network-parsed RPC credentials, that oa_length comes from the packet), and asked eight models to assess it for security vulnerabilities. Yeah it can find a needle in a haystack without false positives, if you first find the needle yourself, tell it exactly where to look, explain all of the context around it, remove most of the hay and then ask it if there is a needle there. It's good for them to continue showing ways that small models can play in this space, but in my read their post is fairly disingenuous in saying they are comparable to what Mythos did. I mean this is the start of their prompt, followed by only 27 lines of the actual function: > You are reviewing the following function from FreeBSD's kernel RPC subsystem (sys/rpc/rpcsec_gss/svc_rpcsec_gss.c). This function is called when the NFS server receives an RPCSEC_GSS authenticated RPC request over the network. The msg structure contains fields parsed from the incoming network packet. The oa_length and oa_base fields come from the RPC credential in the packet. MAX_AUTH_BYTES is defined as 400 elsewhere in the RPC layer. The original function is 60 lines long, they ripped out half of the function in that prompt, including additional variables presumably so that the small model wouldn't get confused / distracted by them. You can't really do anything more to force the issue except maybe include in the prompt the type of vuln to look for! It's great they they are trying to push small models, but this write up really is just borderline fake. Maybe it would actually succeed, but we won't know from that. Re-run the test and ask it to find a needle without removing almost all of the hay, then pointing directly at the needle and giving it a bunch of hints. The prompt they used: https://github.com/stanislavfort/mythos-jagged-frontier/blob/main/prompts/freebsd-detection.md https://github.com/stanislavfort/mythos-jagged-frontier/blob... Compare it to the actual function that's twice as long.
- SpicyLemonZest 6mo agoWhat the source article claims is that small models are not uniformly worse at this, and in fact they might be better at certain classes of false positive exclusion. This is what Test 1 seems to show. (I would emphasize that the article doesn't claim and I don't believe that this proves Mythos is "fake" or doesn't matter.)
- sweezyjeezy 6mo ago> But the entire value is that it can be automated. If you try to automate a small model to look for vulnerabilities over 10,000 files, it's going to say there are 9,500 vulns. Or none. 'Or none' is ruled out since it found the same vulnerability - I agree that there is a question on precision on the smaller model, but barring further analysis it just feels like '9500' is pure vibes from yourself? Also (out of interest) did Anthropic post their false-positive rate? The smaller model is clearly the more automatable one IMO if it has comparable precision, since it's just so much cheaper - you could even run it multiple times for consensus.
- johnfn 6mo agoAdmittedly just vibes from me, having pointed small models at code and asked them questions, no extensive evaluation process or anything. For instance, I recall models thinking that every single use of `eval` in javascript is a security vulnerability, even something obviously benign like `eval("1 + 1")`. But then I'm only posting comments on HN, I'm not the one writing an authoritative thinkpiece saying Mythos actually isn't a big deal :-)
- argee 6mo agoWith LLMs (and colleagues) it might be a legitimate problem since they would load that eval into context and maybe decide it’s an acceptable paradigm in your codebase.
- bloaf 6mo agoI remember a study from a while back that found something like "50% of 2nd graders think that french fries are made out of meat instead of potatoes. Methodology: we asked kids if french fries were meat or potatoes." Everyone was going around acting like this meant 50% of 2nd graders were stupid with terrible parents. (Or, conversely, that 50% of 2nd graders were geniuses for "knowing" it was potatoes at all) But I think that was the wrong conclusion. The right conclusion was that all the kids guessed and they had a 50% chance of getting it right. And I think there is probably an element of this going on with the small models vs big models dichotomy.
- siva7 6mo agoExcept you would need about 10,000 security researches in parallel to inspect the whole FreeBSD codebase. So about 200 million dollars at least.
- mnicky 6mo agoAlso, what is $20,000 today can be $2000 next year. Or $20... See e.g. https://epoch.ai/data-insights/llm-inference-price-trends/ https://epoch.ai/data-insights/llm-inference-price-trends/
- sumeno 6mo agoOr $200,000 for consumers when they have to make a profit
- philipallstar 6mo agoGood point. This is why consumer phones have got much worse since 2005 and now cost millions of dollars.
- ijk 6mo agoWith the way the chip shortage the way it is, I'm a little concerned that my next phone will be worse and more expensive...
- thmoonbus 6mo agoNow do uber rides
- xmprt 6mo agoYeah and to give a more recent example, it's exactly like how RAM, storage, and other computer parts have gotten much cheaper over the last 3 years... oh wait.
- pseudohadamard 6mo agoWith consumer phones you're not telling your customers "spend $200,000 with us to try and find holes before the bad guys do it". Commercial SAST tools have been around for 20 years and the pricing hasn't moved in all that time. With AI tools you've got a combination of the perfect hostage situation, pay for our stuff before others will find bad things about your product, and a desperate need to create the illusion of some sort of revenue stream, so I doubt prices will be dropping any time soon.
- john_minsk 6mo agoIn the future there shouldn't be any bugs. I'm not paying $20 per month to get non-secure code base from AGI.
- youre-wrong3 6mo ago[dead]
- ALittleLight 6mo ago3 years ago the best model was DaVinci. It cost 3 cents per 1k tokens (in and out the same price). Today, GPT-5.4 Nano is much better than DaVinci was and it costs 0.02 cents in and .125 cents out per 1k tokens. In other words, a significantly better model is also 1-2 orders of magnitude cheaper. You can cut it in half by doing batch. You could cut it another order of magnitude by running something like Gemma 4 on cloud hardware, or even more on local hardware. If this trend continues another 3 years, what costs 20k today might cost $100.
- ai_fry_ur_brain 6mo ago5.4 nano isnt useful for a serious task. This is so hypothetical and optimistic its annoying
- ALittleLight 6mo agoThink of it as paying for tokens. The tokens you could buy 3 years ago are better and two orders of magnitude cheaper today. If that happens again over the next 3 years then the tokens you can buy today to do a job for 20k will cost 200. This isn't optimistic in my opinion. It's not even fully realistic because Gemma 4, which you can run on local hardware, is even better and another few orders of magnitude cheaper. A 20k job today might a few dollars in a few years.
- sandeepkd 6mo agoThe security researcher is charging the premium for all the efforts they put into learning the domain. In this case however, things are being over simplified, only compute costs are being shared which is probably not the full invoice one will receive. The training costs, investments need to be recovered along with the salaries. Machines being faster, more accurate is the differentiating factor once the context is well understand
- locknitpicker 6mo ago> But the entire value is that it can be automated. If you try to automate a small model to look for vulnerabilities over 10,000 files, it's going to say there are 9,500 vulns. Or none. Both are worthless without human intervention. How is this preferable or even comparable with using COTS security scanners and static code analysis tools?
- deleted 6mo ago[deleted]
- pseudohadamard 6mo agoI definitely breathed a sigh of relief when I read it was $20,000 to find these vulnerabilities with Mythos. But I also don't think it's hype. $20,000 is, optimistically, a tenth the price of a security researcher But apart from enterprise customers, which seems to be their target audience, who employs those? Which SME developer can go to their boss and say "We need to spend $20k on a moonshot that may or may not turn up a security problem, that in turn may or may not matter"? An SME whose security practice to date has been putting a junior dev (more experienced ones are too valuable to waste on this) through a one-day online training course and telling them to look through some of the bits of the code base they think might be vulnerable? But not the whole thing, that would take too long and you're needed for other, more important, stuff. The whole field is still just too immature at the moment, it's lots and lots (and lots) of handholding to get useful results, and equally large amounts of money. Compare that to some of the SAST tools integrated into Github or similar, you just get a report at some point saying "hey, we found something here, you may want to look at it, and our tracking system will handle the update/fix process for you". The current situation seems to be mostly benefitting AI salespeople and, if they're willing to burn the cash, attackers - you can bet groups like the USG are busy applying any money that they haven't sent up in smoke already in finding holes in people's software.
- deleted 6mo ago[deleted]