3 ms·
Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/ https://cvd.z.ai/ Most of these
by z4y5f3 2mo ago
Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/ https://cvd.z.ai/
Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high.
I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago?
- SyneRyder 2mo ago> ... Anthropic's Project Glasswing is supposed to find them quite a while ago? That was my thought too. For all of Anthropic's talk about their "adversaries", it seems Z.AI have been quietly offering fixes for single shot Remote Code Execution flaws in US software (Safari / WebKit) that Apple and Glasswing / Mythos missed, and that Apple would not attribute to GLM.
- chvid 2mo agoWho says they missed them? Could also be sitting pretty in CIA’s long list of ready to go Vault7-like exploits.
- stingraycharles 2mo ago> That was my thought too. For all of Anthropic's talk about their "adversaries" It’s very likely they found all of them, but that the same happened that happened to Microsoft a couple of decades ago: NSA orders not to disclose / fix them so that they can put it in their collection of unfixed zero days.
- delichon 2mo agoThis is a coherent explanation for why federal model censorship has started with cyber capabilities. But this GLM model release is an in-your-face challenge to that policy. They now have to either set models free or impose a censorship regime that will put anyone not under it at an advantage. Or muddle along in the middle as usual.
- ben_w 2mo ago> Or muddle along in the middle as usual. I'm not a gambling person, but if I was this would be my bet.
- z4y5f3 2mo agoThen "security through secrecy" is really bad mantra especially in the age of AI: others will find the same zero days very soon. If they attack you, then this loses the whole plot. If they propose a fix, then your arsenal becomes smaller.
- tsss 2mo agoProbably Anthropic found them too and promptly got a call from Isreal to stop looking.
- oefrha 2mo ago> and that Apple would not attribute to GLM That was a wtf to me, so I checked Apple’s latest iOS release security content and GLM & z.ai is mentioned once (under WebKit), Anthropic is mentioned twice, Codex is mentioned once. Not clear if there are other instances where the model did most of the work but wasn’t credited. I didn’t bother to check other releases. https://support.apple.com/en-us/128066 https://support.apple.com/en-us/128066
- deleted 2mo ago[deleted]
- re-thc 2mo ago> Anthropic's Project Glasswing is supposed to find them quite a while ago? Someone still has to run it. The analysis and fix could be someone's machine but not committed / published.
- sscaryterry 2mo agoInteresting... So Chinese models are not so bad?
- cromka 2mo agoLooks like they're going for good PR now, to avoid smearing by the "Western" models. Smart!
- croon 2mo agoI'd love to live in a society where people and corporations do good things for PR.
- blooalien 2mo ago> I'd love to live in a society where people and corporations do good things for PR. Maybe so, but I'm not sure I'd like to live in China of all places. (Don't get me wrong. Lotta places I'd like to visit if I ever got the chance, and China's on that list, but to live there? I don't think so.) Maybe one of the Nordic countries?
- vjvjvjvjghv 2mo agoPR for good things doesn’t make money.
- zorked 2mo agoThere's a chance that the real reason why they want to ban Chinese models is that they are so good at fixing bugs and preventing exploits that intelligence agencies have been using for espionage and surveillance for a long time.
- maipen 2mo agoDo you actually believe this?
- 2mo ago
- ThouYS 2mo agoamazing! huge clusters in code from the 1980s haha
- dgellow 2mo ago> and Anthropic's Project Glasswing is supposed to find them quite a while ago? We cannot trust a single company to report security issues, it’s good to see competition in that domain
- ofjcihen 2mo agoOpen source competition no less.
- mcintyre1994 2mo agoIf you look at the distribution of their findings in the linked post, most of theirs are issues introduced a long time ago, almost all before 2006. Complete speculation, but I wonder if they and Anthropic are scanning very different codebases and Anthropic's skew would be in the other direction.
- rbehrends 2mo ago> I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago? You have to consider that having an LLM scan for vulnerabilities is hardly infallible. It is a search guided by heuristics and given a large enough codebase, it is unlikely to identify all vulnerabilities. Personally, I've had Fable 5, GPT 5.6 Sol, and GLM 5.2 all looking for correctness issues in an old abandoned WIP codebase of mine and all of them found some that the others hadn't discovered. Now, correctness issues aren't the same as vulnerabilities, but the same principle about using heuristics to find defects applies.
- Majromax 2mo ago> [A]ll of them found some that the others hadn't discovered. Now, correctness issues aren't the same as vulnerabilities, but the same principle about using heuristics to find defects applies. This makes perfect sense, but that conflicts with the impression put forward by Anthropic and OpenAI (in particular) that they alone occupy 'frontier model' spots. Frontier models should large dominate their competitors on a capability basis, but if GLM 5.2 (now 5.3) is routinely finding bugs / vulnerabilities missed by Fable and Sol then GLM might be genuinely a frontier-grade model by itself.
- Macha 2mo ago“Company hypes own product, downplays competitors” is still a thing with AI
- rbehrends 2mo ago> This makes perfect sense, but that conflicts with the impression put forward by Anthropic and OpenAI (in particular) that they alone occupy 'frontier model' spots. Not necessarily. Even near the frontier, we don't really have a total ordering of capabilities, but a partial order. And even frontier models make plenty of mistakes. Combined with the randomness inherent in searching large codebases for vulnerabilities or correctness issues, it is entirely plausible that even much weaker models (and GLM-5.2 isn't even weak) can stumble upon issues that stronger models missed. My current hypothesis – for which I have only limited evidence, unfortunately – is that it is better to have multiple reasonably powerful (but not necessarily frontier) models looking for issues than just one very powerful one. And even then you're likely to miss out on some issues.
- fsndz 2mo agothis is impressive and actually matches my expectations in terms of near term AI progress. we are going to continue to seem impressive progress in coding & related, anything where verifiability is scalable in an automated way: https://transitions.substack.com/p/a-quantum-of-ai-progress?r=56ql7&utm_campaign=post-expanded-share&utm_medium=web https://transitions.substack.com/p/a-quantum-of-ai-progress?...
- dzonga 2mo agoWordpress having a high number of vulnerabilities not surprising lol
- andai 2mo ago> but isn't the cost for such a scan getting lower by the week Not with Anthropic's models!
- jayd16 2mo agoIn a similar vein, does anyone know how to classify the kinds of problems that are being found? Is it possible to build heavier traditional linting to catch whatever is being caught in a more deterministic way? It seems to me that would be far more efficient in the long run (even if the efficiency is only for the AI to know that aspect was already checked).