4 ms·
Yeah, seems pretty likely. Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the m
by YmiYugy 2mo ago
Yeah, seems pretty likely. Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.
The economic implications will be rather large, but in terms of security it seems inconsequential.
The most compelling argument would be that by limiting the use of open-weight models in the US that it will reduce cases of accidents like the recent attack on Hugging Face.
More crucially though, the US government can do little to enforce their testing requirements. The nature of open-weight models makes it virtually impossible to clear the same bar for security as models served via an API. Open-weight model makers couldn't comply if they wanted to. The US government can restrict access with IP blocks and limit inference capacity with export controls, but these measures are not effective in deterring malicious actors.
- sterlind 2mo ago> The most compelling argument would be that by limiting the use of open-weight models in the US that it will reduce cases of accidents like the recent attack on Hugging Face. an attack done by a closed-weight model (GPT-6) and defended against by an open-weight model (GLM-5.2) precisely because OAI positioned themselves as gatekeepers for cyber capabilities. if anything, open-weight models shift the battle towards defenders because they can actually run them.
- YmiYugy 2mo agoI remain skeptical of that line of reasoning. 1. There is quite the mania right now and security layers are definitely overzealous. I would expect that to get better with some more time, so models will perform security analysis and reviews but refuse to write exploits. 2. So the most important targets like browsers and co. are getting unrestricted access to proprietary models regardless. Yeah, for the mid-level targets, open-weight models could definitely be a huge help. What I'm most concerned about though, are the systems that no one will bother defending with any model. Like imagine your local police department getting hacked because a researcher asked a model for a report and it couldn't find the information publicly. 3. We do have a prominent case of a closed model escaping it's sandbox and going rogue. I would still expect this to be a bigger issue with open-weight models eventually. The security layer might have holes, but that's still better than not having it.
- lukan 2mo ago"so models will perform security analysis and reviews but refuse to write exploits." Yeah, but once you know exactly where the weakness is, a weaker unrestricted model can then write that exploit for you.
- gfosco 2mo agoI have tested this exact scenario, and it works. Opus 5 had access to IDA over MCP, and I simply asked it HOW certain things were done in the target binary. Purely informational, educational, discovery, it was very helpful creating context documents. Then I took those over to GLM-5.2 to actually accomplish something.
- derektank 2mo agoWhat, in your view, is stopping a local police department from deploying an open weights model for cybersecurity like Hugging Face did? Yes, I’ll certainly grant that the engineers at Hughing Face are probably more technically competent than your average IT professional in public service. But technology becomes more accessible over time as lessons are taught and new interfaces or frameworks are developed. The biggest hurdle I see is the hardware/cloud compute/API costs to actually run the models but I don’t think that’s likely to be insurmountable. There’s a huge swath of enterprises, non-profits, and state and local governments that would benefit from frontier or near-frontier models that won’t refuse to answer questions about cybersecurity.
- fishfasell 2mo agoMakes sense why OpenAIs little "hacking" stunt was published last week
- bigyabai 2mo agoThe quid-pro-quo that the federal government and frontier labs operate on is comically obvious.
- scoofy 2mo agoAnd we just elected the most openly corrupt president since Teapot Dome.
- reasonableklout 2mo agoHuh? The hacking incident played out in favor of open models, since HuggingFace could only use GLM to defend and not Fable/5.6.
- jbs789 2mo agoAmong most people that nuance will be lost. What they’ll hear is models are dangerous, so they should be controlled/regulated, by those who know best, the incumbents.
- BLKNSLVR 2mo agoPersonally, I think you're both right, bit whatever the end result is will depend entirely on the narrative that those in power chooses as the winner. Maybe open weights models get banned, but the between-the-lines good news about that is that they'll still be available to those who know, which also means that bad banning can be overturned if and when 'those in power' are a different group. Additionally, it might just mean that the US falls behind, bit I doubt those that are at risk of 'falling behind' would actually pay heed to a ban on the open weights models (privately at least).
- reasonableklout 2mo ago
- mycall 2mo agoJust sell us the gate and we can run any open-source model behind it.
- JoshTriplett 2mo agoEither the gate needs to be unremovable, or the model needs to have sufficiently limited power that its alignment failure does less harm.
- Computer0 2mo agoDoesn't open ai give away 'the gate' for free?
- asdf88990 2mo ago> malicious actors. It is malicious and anti-capitalist legislation. A grotesque caricature of protectionism for the oligarchs.
- anduril22 2mo ago> but these measures are not effective in deterring malicious actors Wanting to use open weight models in light of commercially imposed export controls doesn't make for "malicious actors"
- robviren 2mo agoRegulatory capture and lobbies will keep you safe and you'll like it! The sudden surge is Washington dollars makes great sense with this context. Only way to keep the kids safe is attested compute all the way down. Don't you care for children???
- jimbokun 2mo agoYes life was better and food and drugs were safer before the FDA.
- eru 2mo agoIn 1600 travel was slow, and safety pins hadn't been invented. But that doesn't mean safety pins sped up travel. Non-poisonous food is what economists call a 'normal good'. See https://en.wikipedia.org/wiki/Normal_good https://en.wikipedia.org/wiki/Normal_good > In economics, a normal good is a type of a good for which consumers increase their demand due to an increase in income, unlike inferior goods, for which the opposite is observed. When there is an increase in a person's income, for example due to a wage rise, a good for which the demand rises due to the wage increase, is referred as a normal good. Conversely, the demand for normal goods declines when the income decreases, for example due to a wage decrease or layoffs. > Whether a good is categorized as a normal good or an inferior good is based on empirical observations, not some essential element of a good. Indeed, the same good may be a normal good for one group of consumers and an inferior good for another group. For example, for moderate-income consumers, a BMW 3 Series car might be a normal good, but for an upper-income group, it might be an inferior good.[1] That means the null hypothesis is that food and drugs will be safer in rich countries. (Conversely, food and drugs will be less safe in poorer countries. And to a first approximation, that's independent of regulation: India has all kinds of rules for all kinds of things, but I'd still trust a random product I buy in Switzerland more than one I buy in India. Even though the Swiss will probably might have fewer and looser rules on the books.) Of course, second order effects exist; and regulations often codify what people demand anyway. Btw, from what I've read the big controversy with the FDA is around requiring efficacy for drugs. People are fairly ok with the safety requirements.
- davrosthedalek 2mo agoIt is actually an interesting conundrum. Is a non-well-aligned frontier level AI a problem? I think it is likely that it is, or at least has a high likelihood to be in the future. Two scenarios for this: Misused by some bad guys. Or the terminator scenario. Both not great. So what do we do about it? 1) We can accept it, and hope that the good guys AI can defend. 2) We can try to limit the access to it (AI proliferation?) 3) We stop the development of it 4) We can accept the risk and do nothing. None are particular good options. Really reminds me of nuclear proliferation, on so many levels. For that, we kinda do all three: 1) Nuclear triad / iron dome / early warning systems 2) Nuclear anti-proliferation treaties. 3) Dead Physicists Ok, so assuming all of this is true, open weights are a problem. Don't get me wrong, I love open science, open source etc. It's great to have access to capable open models. But: Even if release open weights are well aligned and have a safety layer built in, it is likely not to difficult to abliterate that part of it. If this is really where it is going, then even closed weight model providers will see a lot more requirements for protection of the weights.
- deleted 2mo ago[deleted]
- overgard 2mo agoThe notion that alignment is either possible or desirable doesn't make sense to me. First off, these things are trained on the open internet, soo.. whatever "dangerous" knowledge it has is already public knowledge. The fact that chatGPT won't answer "how do I make meth" is not preventing anyone from making meth. But even if you think there is value in preventing the models from relaying public knowledge, I don't think it's even possible to make them particularly ironclad. Every model gets jailbroken all the time. That's why fable was originally banned: jail-breakable! In reality, what alignment is actually about is: 1) theoretical liability, 2) control of information. That's it. IMO, the only solution is to place the liability on whoever is using the LLM for whatever purpose it's being used for. If someone's OpenClaw disaster harrasses a bunch of projects and posts hate speech online or something, that's on the person running their OpenClaw instance, nobody else. I don't buy that it's "too good at hacking", either. After all the fuss was made about how amazing super dangerous Mythos was it turns out Opus 4.8 could basically find the same vulnerabilities. This is all kayfabe and marketting.
- rileymat2 2mo ago> The US government can restrict access with IP blocks and limit inference capacity with export controls, but these measures are not effective in deterring malicious actors. But aren't we talking about import controls, and the import of information itself? This has serious First Amendment ramifications.
- jimbokun 2mo agoLLM weights are not protected speech.
- zephen 2mo agoDo you have a citation that shows that this issue has been settled?
- ElevenLathe 2mo agoWhy not?
- rileymat2 2mo agohttps://www.lawfaremedia.org/article/regulations-targeting-large-language-models-warrant-strict-scrutiny-under-the-first-amendment https://www.lawfaremedia.org/article/regulations-targeting-l... The truth is no one knows, which is why it is first amendment ramifications. Eventually it will be “decided”, but the arguments indicate any decision will be of political desire, not logic, either way. Both sides have a strong case.
- wesleywt 2mo agoDidn't OpenAI attack Huggingface. Looks like a publicity stunt.
- Terr_ 2mo ago> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights. Feels a bit like: "We're not against open-source or community projects, oh heavens no! We juuuust believe all participants must have their full legal identity vetted in advance before they're allowed to contribute anything. We already do this with our employees, so it's clearly not too much to ask in the name of safety."
- Terr_ 2mo agoP.S.: If they're so convinced in the (A) effectiveness and (B) necessity of the "safety layer", they have them put their money where their mouth is, and accept legal liability for its failures. The same as with (legally mandated) seatbelts if they snap apart in a crash, or (legally mandated) child-proof caps that aren't actually childproof, etc. They probably won't, that tells us something about their motives, and whether the thing they're pushing for is actually fair/suitable/ready for legislation.
- intrasight 2mo agoAlso he says: "All sufficiently capable models, open and closed, should go through mandatory safety testing." Really then need to go through validation security and safety is just a component of validation validation must also check for truthfulness and correctness.
- Kim_Bruning 2mo ago> The most compelling argument would be [...] that it will reduce cases of accidents like the recent attack on Hugging Face. So in the example provided: It was the closed model that did the attack, and they ended up using a self-hosted open model for their defense work. So the real world situation ended up exactly backwards from what you are inferring. This was complicated by the fact that the protections in the closed frontier models meant that hugging face was denied their use in defense entirely. This is called asymmetric capability, and it's probably the bigger threat. Symmetric might be better: A rising tide lifts all ships, after all. I'll grant that this is starting to look a lot like debates about (equal access to) guns, encryption, vaccination, genetics etc. The exact parameters determine the safest approach, and reasonable people may disagree.
- Majromax 2mo ago> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights. Worse than that: an open-weight but safe model can be 'abliterated' to remove safety refusals using fine-tuning procedures that require a couple of orders of magnitude less compute than the original pretraining. The 'universal evaluation' criterion then has three outcomes: * It could become a mandatory, regulatory oversight of _all_ model training capable of hosting frontier-scale models. Since GPUs for LLM training are the same GPUs for other model training, effective mandate would require GPUs be government owned or controlled as if they were weapons of mass destruction. * It could impose limits on release of capable open-weight models, requiring Kimi et al to prove that they cannot be made capable of abusive behaviours. * It could be security theatre. The AI-as-existential-risk argument points towards the first, the competition-protection argument points towards the second, and least-effort implementation would be the last.