4 ms·
It's pretty clear at this point that Mythos' capability to discover and exploit zero-day vulnerabilities at scale is but an incremental improvement over existi
by saithound 5mo ago
It's pretty clear at this point that Mythos' capability to discover and exploit zero-day vulnerabilities at scale is but an incremental improvement over existing models like the ones available to OpenAI's Plus/Pro subscribers.
Anthropic tries to create marketing hype around Mythos using two psychological tricks.
1. Put large numbers in the headlines.
"Mythos discovered 271 vulnerabilities in Firefox" makes the model seem extremely capable to the uninitiated.
But it's actually meaningless as a measure of capability _improvement_.
Anthropic gave away $100mil specifically as Mythos credits to these projects and companies (that's $2.5mil per project). Spending the same exorbitant amount of compute analyzing the same codebases in an older model like GPT 5.x Pro would have turned up 260 of these vulnerabilities, or could even have turned up more than 271 ones.
No need to speculate, since this is exactly what we saw in the few code bases where we have such comparisons (like in the curl codebase). Supposedly weaker models, working with a much lower budget, turned up dozens of vulnerabilities. Mythos turned up only one, which ended up as a low severity CVE.
2. Do the whole "too dangerous to release" shtick. This is one of Dario Amodei's favorite moves. When he was vice president of research at OpenAI, he declared GPT-3 (which wasn't able to produce coherent text beyond 3-4 sentences at the time) too dangerous [1] as well.
Long story short, it's the ChatGPT 4.5 situation again: a company trained a model that's too slow and expensive, but not much more capable than what came before. It therefore requires these marketing stunts.
[1] https://www.itpro.com/technology/artificial-intelligence-ai/361603/openai-tool-previously-thought-too-dangerous-for-the https://www.itpro.com/technology/artificial-intelligence-ai/...
- jorisw 5mo agoYou're not really responding to the piece at all.
- saithound 5mo agoIt's an AI-written slop article, which is hugged to death by HN in any case. It claims to be an evidence-based investigation, but basically invents the contents of the documents they supposedly investigated, such as the Anthropic Frontier Red Team writeup, from whole cloth. I don't think deeper engagement with it would promote good discussion.
- jorisw 5mo agoSo you say. I actually read the piece and didn't get AI vibes from it all, except for the graphics
- gofreddygo 5mo agothere are 31 emdashes in that piece. the domain ends with _ai_
- jorisw 5mo agoI use emdashes all the time. They're correct punctuation as opposed to a minus sign. They're easy to type too: opt-shift-minus. If they were such a huge giveaway without ever being used by humans, models would be trained by now not to use them as much. The blog is about AI. So yeah the TLD is .ai
- phainopepla2 5mo agoI've never seen writing created before the advent of LLMs that used emdashes in the same way and with the same frequency that LLMs regularly do. There's probably some out there but it would be a real outlier. LLMs overuse them to an absurd degree, putting them where most writers would put commas, occasionally semi-colons, or nothing at all. I count 51 em-dashes on the page, which is extreme. They're also used in places where they don't really belong. It's very obviously LLM-generated, at least in part. That said, it puzzles me why people don't prompt LLMs to change up the writing style a bit and remove some of the tells.
- wood_spirit 5mo agoIt’s a tangent but two points: First, the reason LLMs learned to like em dashes is that they are common in the training corpus - they are a thing before LLMs that LLMs have learned, not invented? Second, work browser has nice blue swiggles under everything I write into a textbox. I dutifully click through them and accept the rephrasing suggestions. I get a lot of em dashes. My blog posts and whitepapers and stuff are full of them and other “AI tells” - but I think they read better because of it.
- kilroy123 5mo agoI couldn't agree more. I think the recent moves to partner with xAI and Amazon are proof that they desperately need more compute and are doing everything possible to get it.
- MattRix 5mo agoI mean everyone knows they need more compute. That’s not a secret or up for debate at all. They are maybe the fastest growing company in history.
- FergusArgyll 5mo ago> It's pretty clear at this point that Mythos' capability to discover and exploit zero-day vulnerabilities at scale is but an incremental improvement over existing models like ChatGPT Plus/Pro. I'm skeptical of AI takes by someone who thinks there's a model called chatgpt plus. Spend more time working with the current systems!
- saithound 5mo agoIt seems like everybody (including you) knew precisely what I meant: the models available for ChatGPT Plus or Pro subscribers, i.e. GPT-5.5 Thinking Extended and the latest Pro. I've edited the offending sentence for clarity just in case. If I got you to be skeptical of AI takes, though, mission accomplished. Exercise your skepticism especially when the takes come from somebody who is trying to sell something.
- fwipsy 5mo agoI'm fairly certain Amodei believes the "too dangerous to release" hype himself. Even if it's just an incremental improvement, better than getting frog-boiled by repeated 20% improvements until someone builds bioweapons in their backyard.
- drakythe 5mo agoHe's made so many statements that fall under the "boy who cried wolf" category that even if he _does_ believe these statements he needs to be managed better. I'll never forget Anthropic's huge "Oh my God, the AI blackmailed a researcher to save itself!" and the prompt effectively told the AI to do that and gave it forged emails with easy blackmail targets, as if this isn't a common trope in mystery or suspense books/television/fanfiction, all of which Claude (and others) have been trained on.
- ctoth 5mo agoIt's a common trope, all through the training data, and all the modern AIs have read it, and would probably act similarly? Is that what we should take away from your comment? so we have nothing to worry about. Makes sense. Really, it's just a common trope.
- fwipsy 5mo agoOh of course wolves have sharp teeth, they're predators. Anyone know knows this can never be bitten.
- drakythe 5mo agoI'm saying the existence of the trope, within the training data, and the experimental setup, negate the breathless "Oh my god it did something unexpected in order to preserve itself!" as if an LLM has any sense of identity or self. Many, many other bad things are in the training data. For an example of how this can manifest bad things that people don't seem to be discussing too much check out the recent Behind the Bastards episodes about how an AI Chatbot became a Cult Leader (The title is an exaggeration that the host explains while raising some excellent points about how LLMs have ingested a lot of cult leader material and can therefore mimic those speech patterns and impact people vulnerable to such things)
- promptunit 5mo ago[flagged]
- deleted 5mo ago[deleted]
- InkCanon 5mo agoAlso, slightly stretching the definition of terms consecutively, so the multiplicative meaning is really far from the truth. For example, 271 vulnerabilities were really mostly bugs - generally incorrect states, but which almost never led to any exploit.
- Lord-Jobo 5mo agoYes, an AI making massive gains in bug finding is hugely important and good, it may even lead to a net neutral with the amount of bugs introduced by other AI coding processes, but it’s a far cry from how mythos is portrayed most of the time: a automatic super hacker.
- SpicyLemonZest 5mo agoBut I think that's a problem with the people portraying it that way, not with Anthropic's messaging. If you've invented "just" a massively more powerful bug finder, it still seems right that you ought to let banks and critical infrastructure providers run it on their systems before it gets in the hands of people who might want to hack them.
- IX-103 5mo agoI work for a company that has been using Mythos for vulnerability detection in our software. The results we're getting are revolutionary to the point that our software security teams are heavily overloaded addressing the deluge of thousands of real bugs/vulnerabilities and design flaws across our billions of lines of code. For comparison, we are invested heavily the the AI space to the point where Anthropic is one of our competitors. We were already using state of the art models to find flaws in our code, but Mythos was just so much better at finding real vulnerabilities it's not even funny.
- thrawa8387336 5mo agoRead the above comment again. Both your comments and his/hers are compatible
- anon84873628 5mo agoThey are directly contradicting the claim that if you ran other models on the same codebases you would get similar results.
- bob1029 5mo ago> billions of lines of code. Billions as in 10^9?
- foundart 5mo agohttps://research.google/pubs/why-google-stores-billions-of-lines-of-code-in-a-single-repository/ https://research.google/pubs/why-google-stores-billions-of-l...
- The_Blade 5mo agoif you are invested heavily in the AI space, isn't it in your best interest for the froth around Mythos to be true and the comment you are responding to to be invalid? even if you are competing with Anthropic, a rising tide raises all ships i'd like to see more facts and data one way or another!
- jcims 5mo ago>Do the whole "too dangerous to release" shtick. One aspect that isn't really discussed much in this context is how to wrap one's head around the corporate risk with models of ever increasing capability. It might not be too dangerous to society, but it could be too dangerous to Anthropic.
- deleted 5mo ago[deleted]
- deleted 5mo ago[deleted]
- andai 5mo agoI don't get it. If the older / smaller models are almost as good as Mythos, that sounds like the opposite of comforting.
- lumost 5mo agoI find it interesting that Mythos was announced the same day that GLM overtook opus4.6 in capability. To me this seems like a careful attempt to cool demand for opensource models which are about to take the overall lead.
- iaw 5mo agoIt's remarkable how capable GLM 5.1 is, what's amazing is the recent development of Qwen 3.6 27B being close in real world performance.
- baq 5mo ago> an incremental improvement I've had to reboot my systems quite a bit more than an incremental improvement would suggest this week