6 ms·
From https://huggingface.co/blog/security-incident-july-2026 https://huggingface.co/blog/security-incident-july-2026 , this is frickin' hilarious: > When we st
by foo12bar 3mo ago
From https://huggingface.co/blog/security-incident-july-2026 https://huggingface.co/blog/security-incident-july-2026 , this is frickin' hilarious:
> When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.
- iugtmkbdfil834 3mo agoIt is pretty funny, because there is something here for everyone. People who don't believe in guardrails have a clear indicator as to why operators should have access to models that don't try to question their Daves. On the other, people, who think that if only we could align the models just a tiny lil bit better, none of this would have happened to begin with. Pure madness.
- foo12bar 3mo agoAnd the fact they used a Chinese model, because none of the frontier models from very highly valuated top US companies support their very common and essential use case.
- jonplackett 2mo agoThere’s an article from yesterday I think it was stratchery where they say it’s also because the Chinese open source models are better because they don’t have to play by the no-distilling rules that the western models have to honour.
- larodi 2mo agoAnd no similar sized big model is even public, which is super important to note.
- ethin 2mo agoI also thought this was hilarious. The very "safety" measures these models have prevented them from doing... Something that is designed to increase safety?
- sillysaurusx 3mo agoIf anyone's looking to actually run a model that doesn't have guardrails, there's an uncensored model that you can run locally with llama.cpp: https://www.reddit.com/r/LocalLLaMA/comments/1rq7jtm/qwen3535ba3b_uncensored_aggressive_gguf_release/ https://www.reddit.com/r/LocalLLaMA/comments/1rq7jtm/qwen353... Specifically, I serve the model with this shell script on my M2 Max: https://github.com/shawwn/scrap/blob/master/llama-serve https://github.com/shawwn/scrap/blob/master/llama-serve It's pretty good. I used it to do some pesticide research. (Normal models all refuse due to guardrails about bioweapons.)
- CamperBob2 2mo agoAlso the HauHau abliteration (uncredited Heretic treatment) of 3.6 27B is excellent, for tasks that benefit from more world knowledge.
- iugtmkbdfil834 2mo agoHeretic truly is the unsung hero. Also, noted HauHau for testing.
- baq 2mo ago> It's pretty good. I used it to do some pesticide research. (Normal models all refuse due to guardrails about bioweapons.) As someone who has never once had any need whatsoever to research pesticides I… don’t think it’s bad at all? I don’t want anyone to have the capability to invent a human-targeted pesticide who isn’t verified not crazy?
- deleted 2mo ago[deleted]
- jrs100000 2mo agoLlama isn't going to invent shit. It wont be able to tell you anything accurate that you couldn't get out of a chemistry textbook.
- samplifier 2mo ago"Dave" seems to be a reference to "2001: A Space Odyssey" where the AI becomes ... cheeky ... and no, not in a Pygmalion kind of way (that's coming soon).
- davrosthedalek 2mo agoI am somewhat more worried about a Darkstar AI.
- friendzis 2mo ago"if only we could align the models just a tiny lil bit better" is a rehashed "if only we could escape untrusted inputs just a tiny lil bit better" from 2000s, that were RIPE with various form of malicious injection. Every command+data channel in existence has been and will continue to be exploited one way or another, because the solution space is for all intents and purposes unbounded. Sure, highly defensive escaping reduces attack surface dramatically, but e.g. prepared statements eliminate the whole class of bugs. As far as I understand, current LLMs are architecturally incapable of this separation. Given the inherently recursive nature of GenAI, the model itself is part of the input space, making validation essentially impossible.
- insanitybit 2mo agoEscaping inputs is at least somewhat tractable. It's unclear if alignment is.
- iugtmkbdfil834 2mo agoBut see... this is why it is a perfect long term job in the age of AI;p Them peoples think they found ultimate hack.
- friendzis 2mo agoNo, that's a misconception that led to decades of vulnerabilities. You see, you can (usually) easily tell what a particular escaping transformation does. That does not tell you neither how it will be interpreted down the line, nor what should be done. Arguably the most common problem is double escaping. This typically manifests as various double escaping bugs. If your hand rolled implementation just chains `.replaceAll("<sepcial>", input)` and `.replaceAll("<escape>", input)` the escaping result depends on evaluation order, at least on already pre-escaped inputs if your instruction sequences are single-element. Even if you get it right and don't reduce escaped sequences to double escape + unescaped, you are still dealing with stray escape sequences: "John o\'Doe". You have to meticulously (all the way through your call tree and even through persistent storage cycle!) track if a value has already been escaped and whether it needs escaping. As paradoxical as it may sound, meticulous and defensive escaping produces tons of bugs. If the project decides to escape raw inputs right when passed in, escaping them once more before passing out (to storage, another process) is a bug that you cannot easily statically test against. On top of that, various modules that you interact with (storage, libraries, modules pulled from another team) will have different escaping semantics: some will apply escaping on their own, some will expect input to be "sanitized" and treat it raw. The semantics can even be different on different paths: write to storage module accepts input as is, but retrieval method "helpfully" runs escaping. Furthermore, in different contexts the escaping rules are going to be different. In a web world, what's safe to write directly to html, pass to js `alert`, and pass to sql query are entirely different things. Input safe to dump into html is not necessarily safe to pass to string-interpolating SQL DTO layer, and vice versa. Then the DBA changes config to allow variables in queries and your escaping _semantics_ are now entirely different. It's a minefield with essentially unbounded surface. You will trip up. People have tried to solve the problem for decades. Very smart people have tried. They all have failed.
- baq 2mo ago...and we have an US company defending itself against an overwhelming cyberattack from another US company using Chinese tech. what a time to be alive.
- Lerc 2mo ago>why operators should have access to models that don't try to question their Daves. I am unsure if this is terminology I am unfamiliar with, a typo of Devs, or a 2001 reference.
- sethammons 2mo agoI immediately took it to be a slick 2001 reference. Devs tell computers what to do. Computers tell Daves "I can't do that."
- biztos 2mo agoIn this new age of AI, all the devs are Daves.
- Sharlin 2mo agoBoth of those groups of people are crazy.
- DubiousPusher 2mo agoYeah, this is very much one of those stories where people from many different perspectives or chopping it up on a plate and ripping it through a straw.
- chinathrow 2mo agoTurtles all the way down.
- runtime_lens 2mo ago[flagged]
- gertrunde 2mo ago> This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment. Well, that may be correct for the second, local, analysis attempt... but seems funny to tout this as an advantage after already having tried the opposite...
- zeroday404 2mo ago[flagged]
- redleader55 2mo agoIt's even funnier because an attack, until proven otherwise, should make you assume the data has already left the environment.
- davrosthedalek 2mo agoWell, I think it's fair to assume that a) They didn't upload everything before they realized it would work. b) They want to mention this as an advantage for future analyses c) Even if you assume that the attack exfiltrated everything until proved otherwise, you shouldn't just disseminate all the private information, because maybe the attack didn't.
- horsawlarway 2mo agoThey specifically call out credentials used during the attack. But they should be rotating those regardless. You don't get to say "Maybe the attacker didn't get this credential". You just rotate. The most generous interpretation is that they have not yet have completed that rotation, and they didn't want to risk putting those credentials into the wild during that process. --- But all of that aside, I feel like the undercurrent of this comment is that the "safety" rules that providers are pushing are genuinely harmful. Another point where "if you don't own the model, you can't properly operate the tool" becomes true. Open isn't about profits, it's about capabilities.
- atcon 2mo agoAgainst an "agentic attack" and compromised credentials, one should be paranoid about latent vulnerabilities [1] [1] eg "Robin Hood and Friar Tuck", poisoned compiler, etc. https://news.ycombinator.com/item?id=26553390 https://news.ycombinator.com/item?id=26553390
- diabllicseagull 2mo agoso an on-premise and open-weight model was more useful than a commercial frontier model?
- djmips 2mo agofor hacking I assume that would be the case
- mfarhanshams 2mo agoyes for an out of syllabus thing