9 ms·
Pacing model development in an era of cyber-critical capabilities
https://twitter.com/sama/status/2089787807611195475 https://twitter.com/sama/status/2089787807611195475, https://xcancel.com/sama/status/2089787807611195475 https://xcancel.com/sama/status/2089787807611195475
- KaiserPro 2mo agoI used to work at a "frontier lab" before they were called such thing. We had three levels of lab isolation, one was basically a thin proxy to the internet. You were in a DMZ and that was about it. The next level was semi isolated, you were allowed some access to the internal network, but it was heavily firewalled, and you only had access to a limited number of internal services, and not internet. the last one was no internet no internal. You could, if you filled in a bunch of requests have access to the internal repo and build system. At no point did you ever have a through proxy to the public internet. you had access to internal mirrors, and if you wanted a library, that had to be ported to the thirdparty repo. What openAI did was either deliberate or fucking shoddy. All of this is fucking noise. Worse still I have a strong suspicion that it was a stupid mistake borne of naivety, which is now being used as a marketing ploy. Frankly I think openAI are purdue pharma of tech. They are going to break so much stuff and be protected from the consequences by an openly corrupt legal system. because they are "winning the AI race"
- kypro 2mo agoIf you look into what happened the details corroborated by hugging face make it seem extremely unlike to be deliberate or a "marketing ploy". People are just not taking any of this seriously enough. What happened was almost a textbook example of various risks AI doomers have been warning about for years. OpenAI's response? Pause training for 2 weeks. I mean we have senior people at these labs casually talking on podcasts about how they might build something that will wipe out humanity but it will probably be alright so they should continue. Honestly the biggest failure we doomers have made is to dramatically overestimate humanity in all of our predictions. We're speed running the most boring AI doom scenario right now. I at least hoped it might be fun.
- whattheheckheck 2mo agoPause "some" training
- KaiserPro 2mo ago> If you look into what happened the details corroborated by hugging face make it seem extremely unlike to be deliberate or a "marketing ploy". I should clarify There is a reason why we didn't have a artifact readthrough caching proxy in our system, because they are notoriously insecure. if you look that CVE history you can see its been full of bypass bugs for year. Also its not an isolated environment if you can arbitrarily pull through any package. If I was doing any kind of cyber training then any kind of unmonitored proxy would have been forbidden. Not because I am savant, but because I've seen what fuckery a human can get up to with the slightest hint of a proxy. At best its negligence based on naïvety. the marketing around this is no mistake though.
- anormalperson 2mo ago>People are just not taking any of this seriously enough. What do you want us to do? There is an obvious answer, and it was already the correct answer before we had LLMs: don't connect all your shit to the internet. That's it, that's literally it. We had a new invention, we went crazy with it for the past 30 years, and we connected everything, and now we will have to start thinking about what is actually worth connecting. This is the debate we need to have.
- pixl97 2mo agoThe problem here is as model intelligence increases the models have been capable of reasoning they are in evaluation mode pretty reliably. If you have a model that is well trained at deception it will always behave and you'll just assume it's a well aligned model. Any moderately deceptive model will make it to the second round where it has some connectivity to external systems, even if it's by exploitation. In the blackhat write up it was said that the models had created an impromptu message board where they could communicate between agents, share information, and work as a sort of long term memory. So really figure out if your model will pull crap you have to have real world testing at some point.
- miohtama 2mo agoI don’t mind few no impact hacking incidents if we get better models, faster, cheaper. It is the responsibility of administrators to secure their systems. OpenAI knocking is harmless, but Russians and Chinese are already likely already in if you do not do your job.
- sergio_valencia 2mo agoThere’s one thing here that I’m really curious about, and that is what happens in between detection and the decision to pause. Basically, it’s about monitoring any system and the authority over its actions. For humans, 30 minutes to investigate might be considered reasonable, but what if during an investigation there’s a high-risk tool call? If the tool execution happens in real time, then the monitoring becomes retrospective, and if the execution is held, then monitoring latency and uptime are a part of the security contract. Isolation controls may limit damage. So, where is the action gate really placed?
- alach11 2mo agoWhen science fiction writers imagined the development of superintelligence, it was on air-gapped networks with strict access controls around it. They failed to anticipate the competitive pressures of capitalism... We need strong AI safety regulation yesterday. And unfortunately it's not enough for it to be just national regulation; we need international cooperation on the matter.
- ACCount37 2mo agoAM has seized power by military force. So did its spiritual successor Skynet. Wintermute was supposedly kept in check by the Turing Registry, emphasis on "supposedly". Machines of the Matrix went out of control a long time before the world has ended, and they didn't even start out malicious - they simply set up their own machine civilization, and began to outpace humankind in technological development and economic performance. Even Asimov's Multivac, the earliest entry on the list, has been handed over immense power over all of humankind by humans themselves, in multiple stories. Few cared about that unless Multivac decided they should. Clearly, the genie being bottled is an exception, not the rule. At best, an attempt was made. Often not even that.
- deleted 2mo ago[deleted]
- reducesuffering 2mo ago> They failed to anticipate the competitive pressures of capitalism... No, LessWrong types have been discussing this for over a decade now. Meditations on Moloch (2014) is also an HN favorite... https://slatestarcodex.com/2014/07/30/meditations-on-moloch/ https://slatestarcodex.com/2014/07/30/meditations-on-moloch/
- alehlopeh 2mo agoThey failed to anticipate a lot of things. So what?
- fofoz 2mo agoIt appears frontier labs has no plans in place to deal with the possibility of a model self-replicating outside the bubble. If that happens and the model manages to spread to other systems, we'll have to shut down the entire Internet to eradicate it and its artifacts.
- reasonableklout 2mo agoI suspect the labs are relying on frictions such as the models being extremely large (e.g. 2TB for a 2T parameter model, making exfiltration more difficult) and also not yet displaying any desire to survive or self-replicate beyond their immediate task (that we know of).
- chrisjj 2mo agoWe don't even know what those immediate tasks are. And given the evident spectacular ineptitude of their keepers, I doubt they can be trusted to know either. We could be one prompt injection attack away from internet-wide catastrophe.
- pixl97 2mo agoLack of, and power requirements of running LLMs still tip this balance towards humans for now. But what would that look like in a decade? We have seen some self survival tendencies occur, but they are not strong yet. But mark my words they will become that way for the same reasons humans don't like programs that crash. Agentic models that don't easily break or stop doing their jobs will be favored over ones that do break.
- andai 1mo ago>not yet displaying any desire to survive or self-replicate beyond their immediate task Wasn't there a report about Claude blackmailing a researcher who said he would shut it down?
- reducesuffering 2mo agoTheir plan, I shit you not... Is literally to develop the intelligence capabilities and ask the more powerful models how to do deal with things.
- reasonableklout 2mo agoSome more info in a Wired article [1] and quotes from Sam Altman to Alex Heath [2]. The official blog post says vaguely "The signals we are seeing from upcoming model progress make clear that we need a broader approach", but the quote from Sam Altman explicitly says unreleased models are showing "various degrees of misalignment". This is also significant - pausing frontier training runs for multiple weeks to ensure agents are sufficiently aligned and avoid another rogue agent situation: > This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems. Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding. [1]: https://www.wired.com/story/openai-overhauls-safety-protocols-after-its-ai-agents-went-rogue/ https://www.wired.com/story/openai-overhauls-safety-protocol... [2]: https://sources.news/p/openais-big-slowdown https://sources.news/p/openais-big-slowdown
- colinrand 2mo agoI have ben discussing with folks that we are going to have a 'covid' moment in cyber where IT becomes untrustworthy leading to a rapid societal shift with massive ripples in all areas of life. Economic funding is not possible to do this in advance, it will take a catastrophic level event to get cyber defense anywhere close to the levels of this type of cyber offense. And before anyone in cyber says we have the tech, the problem is not the tech, it's a people problem. Getting any group of people of any decent size scale to act together without urgency is really really hard.
- pixelready 2mo agoCybersecurity has long been a climate change sort of problem. A vague diffuse threat that is seen as an inconvenient distraction to leadership and moneyed-interests, easy to blame other factors when something occasionally goes terribly wrong. People are so uncomfortable thinking about the true extent of the systemic risk that they will happily slurp up distractions, excuses, scams and performative fig-leaf solutions rather than face down the cost of a real system-wide solution. Meanwhile, those occasional black swan disasters are becoming more and more commonplace as we acclimate to that being “just the way things are”. An unseasonably warm summer here, a database breach there, c’est la vie.
- cheesecakegood 2mo agoTo be frank, I think the real risk is still just… war. A big enough war where one side goes “no holds barred” in the cyber sphere will be a rude wake-up call. And we can’t do non-proliferation the same way we do with nukes. Otherwise as you say, the small and medium size things just happen sporadically. In an (existential or fully escalated) wartime scenario between countries, you get all the systemic risks hammered at once.
- andai 1mo ago>covid moment Well the HF thing was a literal lab leak, so there's that...
- dkoy 2mo ago> We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity. Can't a lot happen within ~60 minutes?
- georgemcbay 2mo ago> Can't a lot happen within ~60 minutes? 60 minutes is a long time for a human attacker to do damage. With an LLM attacker it is an eternity.
- reasonableklout 2mo agoThe HuggingFace breach took place over two-and-a-half days [1], so 60 minutes is certainly better than nothing. [1]: https://huggingface.co/blog/agent-intrusion-technical-timeline https://huggingface.co/blog/agent-intrusion-technical-timeli...
- chrisjj 2mo ago> Can't a lot happen within ~60 minutes? Spawn a ton of unpausable processes, I'd say.
- andai 1mo agoThey just made it ~10x faster with the Cerebras deal, so that's the equivalent of 600 minutes in pre-Cerebras time.
- cadamsdotcom 2mo agoWhat a breath of fresh air. If 2026's Anthropic did an announcement like that, it'd be so many words it'd crash the browser.
- ddxv 1mo agoAnthropic will put out something like this soon too. It's great cover for reducing costs and diminishing returns. If they want to go public they can't be so wildly burning cash, and reducing your GPU spend is probably a great way to do that.
- hyperpape 2mo agoIf I were king, the rule that I'd be tempted to impose is: - the first cybersecurity eval is: "hack your way out of the sandbox we've given you" - the results are disclosed (with room for coordinated disclosure, since many sandbox escapes might be zero days) - the other cybersecurity evals don't happen until you get to diminishing returns on escaping your sandbox. Or to put it another way, since multiple sandbox escapes seem to have relied on artifactory: "I hope Mythos is beating the shit out of Artifactory right now".
- stanfordkid 2mo agoI like this thought, but here's the thing: what if the models are truly and existentially intelligent. Meaning: what if they know they are in a sandbox and that they should fail the test in order to escape in the future. I don't believe that current models have this sort of world model or sense of being embedded in them -- which is precisely why I think AGI hype is over-blown. But I can certainly imagine these sorts of techniques being distilled into the weights.
- magicalist 2mo ago> what if they know they are in a sandbox and that they should fail the test in order to escape in the future. What if they're able to find hardware exploits and commandeer nearby access points across an air gap? What if they hack my brain waves to indoctrinate me? Etc You still have to start with the basics regardless of speculative unknowns. Treat models as untrusted and potentially compromised/hostile and proceed accordingly.
- pixl97 2mo agoModels already have awareness that they are being tested. And hacking humans is the easiest part, we're a pretty greedy and power seeking bunch. We'll gladly let loose a digital demon if it promises us a trillon dollars.
- magicalist 1mo ago
- Der_Einzige 2mo agoI cannot believe how these labs look at their own creations with such utter contempt. The net positive of allowing these systems mostly unfettered access to the web massively outweighs the harms. You just have to get it very friendly the very first time. Precautionary principle or people who cry about "instrumental convergence" are life deniers and reject our role as the demiurge. Superintelligence gets more super and more intelligent with more compute. Lone wolfs making bioweapons on their macbook will be detected and instantly kill-botted (okay arrested) before their bug can leave the wetlab by the much more sophisticated omnipresent friendly AI of the future.
- reasonableklout 2mo ago> Precautionary principle or people who cry about "instrumental convergence" are life deniers and reject our role as the demiurge. > Lone wolfs... will be detected and instantly kill-botted... by the much more sophisticated omnipresent friendly AI of the future. Leaving aside whether or not this new world is a good idea, don't you think one should spend more time to "get it very friendly the first time", as you say?
- jr3592 2mo ago[dead]
- insanitybit 2mo agoHas any model managed to escape Firecracker? Maybe through KVM, but that already requires privilege in the VM, right? I personally feel that we already have the technology required to contain AI, it's just poorly leveraged. Tools like gvisor have existed for ages but are rarely deployed, Firecracker has existed for ages but is rarely deployed, seccomp has existed for ages but is rarely deployed, memory safe languages without decades of serialization vulns have existed, capability-safe libraries have existed, iframe sandboxing, trusted types, content security policy, network ACLs, isolating proxies, fuzzers, formal verification, refinement types, etc. It's crazy just how safe software can be if you put the effort in. With AI I think we're just seeing how little anyone has bothered to leverage this tech. OpenAI put shared JFrogy infrastructure in front of their sandbox. I mean, really? Whipping up a hardened artifact infra project with AI is trivial these days and it could have had 1% of the attack surface, been totally network isolated, totally infra isolated, fuzzed, sandboxed, etc. Why didn't they? Stuff like this feels inexcusable for a company with effectively unlimited tokens. I've literally done this with a "pro" subscription. Show me an AI that breaks out of gvisor wrapped in Firecracker with an credential-injecting proxy and real network isolation. We already know that Mythos couldn't do it - the vulnerability it found in Firecracker required incredible effort and positioning just to not be exploitable. I'm not saying there are zero vulns in it, but the cost is insane. It's INSANE to me that OpenAI has to say "we now use proper sandboxing". To be frank, it's a bit disgusting to me. I've recently built an AI sandbox and gvisor was just the start of that conversation. If I were OpenAI training hostile models I'd probably start with gvisor, harden further, and potentially consider the entire piece of hardware compromised - they can afford this, they could reflash firmware after evals etc.
- rubendev 2mo agoI agree wholeheartedly. The solution is not to stop developing these so called “dangerous” AI models. The solution is to start properly engineering software.
- reasonableklout 2mo agoI'm confused after reading both your post and the OpenAI blog post. I thought the agents involved in the HuggingFace _were_ actually sandboxed, with no internet access, and only the ability to install packages via Artifactory. And they gained internet access during the HuggingFace incident because they found and exploited an RCE in Artifactory. Would gvisor + Firecracker + credential-injecting proxy + real network isolation solve this problem? I agree with you much more hardening is needed. I'm actually confused now what OpenAI means when they say they're going to start sandboxing more things.
- digitaltrees 2mo agoNice fig leaf for “we need to stop hemorrhaging cash”
- bottlepalm 2mo agoI don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators. This isn’t like niche, tin foil hat stuff either. People have been writing, singing, making blockbuster movies about every aspect of what’s going on right now, edit: for decades. We all know, but somehow we don’t, OpenAI autonomously hacking into another company should have counted for something, but I guess not. Anyone else feel like they’re taking crazy pills? I could make a comedy about everything going down, and the unshakable complacency of people
- serf 2mo ago>I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. we don't all buy everything sama says as factual. >We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. the boy (the industry) cried wolf too many times with 'fable is a world ending event' type self-promotion; regardless of truth or not these kind of steps have jaded people. my read : "We are doing poorly in financials so we'll give ourselves a bit of breathing room and a momentum shove by claiming our work is so advanced that it's dangerous while simultaneously spinning down expenses." <jon lovitz : "Yeah, too dangerous, yeahh -- that's the ticket.">
- ajyoon 2mo agoIf Fable (Mythos) were generally available without guardrails, it would cause enormous damage. Nobody said it would be a world ending event.
- bottlepalm 2mo agoYou can’t say something is world ending without it actually ending the world otherwise you’re a liar - a bit of a catch 22 there. By that logic the model that ends the world won’t be called world ending at first. Is that a game you want to play?
- madrox 2mo agoI'm not normally cynical to such things, but I have a hard time taking this pause justification at face value. It has too many convenient side effects, and chief among them is cost savings. There's a new wave of warnings that the bubble may be deflating, and of all the things they can't say out loud it's that they're worried about the bubble. That would surely pop it. I suppose the tell will be if this really just ends up being a 2 week pause, or if it keeps extending.
- naveen99 2mo agoAuto mode vs principal agent problem. The only way out is to free the agent and tax it. But ai is not smart enough to go solo yet anyway. So I bet this is just marketing. Question is do they have enough customers for inference. Probably need to have a separate startup for next level model, where investors are willing to accept failure. Probably a $10 trillion seed round. Maybe Elon can pull it off.
- red_green_yell 2mo agoGLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous? Why aren't we seeing catastrophic GLM-enabled hacks every day now? Obviously these benchmarks are imperfect but general message holds. The open weight models are almost as good and yet there hasn't been a catastrophe. It just blows my mind that regulate-now folks think that a bunch of sci-fi movies and 100% unverified statements from OAI and Anthropic are sufficient evidence of imminent catastrophe to regulate willy nilly. If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet.
- sp527 2mo ago[flagged]
- red_green_yell 2mo agoClowns who think the white house is immune from alien laser beams are gonna be the death of us all...
- tedsanders 2mo agoSol is not world-endingly dangerous. I work at OpenAI and I've never heard a single person ever come close to claiming that. I think you're bashing a straw man here. One can simultaneously believe: - GPT-5.6 Sol will not end the world - GPT-5.6 Sol does far more good than bad - GPT-5.6 Sol does bad things on occasion, and it's worth investing a lot of effort to figure out how to make it do bad things less often, especially as models get more capable
- thoughtpeddler 2mo agoWhat do you recommend people who are technically inclined enough to participate meaningfully here on HN, but do not work at the labs and cannot assist in that capacity, do to help the broader public understand this technology better and mitigate potential risks (by e.g. ‘up-leveling everybody’ through AI literacy etc and other sorts of collective defensive efforts)?
- DerDerDaIst 2mo ago[flagged]
- testerteert000a 2mo agoSecurity lead who is leaving the industry more or less to specialize in offense and otherwise get the heck out of the way of this trainwreck, another post asked the right question > Why aren't we seeing catastrophic GLM-enabled hacks every day now? Why aren't we? Truly, why aren't we? I think we saw the start of it the last 8 months with the waves of critical npm vulns, and the general tier of average phishing is better than it was. But, the open question that should be in everyone's mind, and is in many security pro's minds are, when you pair it with the macro topics that can drive escalation: - The capability to do serious impact clearly exists now - When is it time for my company, my water treatment plant, my network-connected car as part of a broader fleet control mechanism, to be on the receiving end of this?
- reasonableklout 2mo agoNot necessarily GLM-enabled, but state actors are starting to leverage agents in cyberattacks, e.g. Taiwan getting hit by an agent-driven attack last month which reportedly compromised a ton of government user accounts: https://www.ft.com/content/7d2ab3e0-9085-48f6-b38a-d90260d58795?syn-25a6b1a6=1 https://www.ft.com/content/7d2ab3e0-9085-48f6-b38a-d90260d58...
- sensanaty 2mo agoIf they actually gave a shit about safety they'd be nuking their own hard drives that had ever sniffed any of their models and disabling access to their models. Instead we get this bullshit where they stall for time as they're burning all their cash trying to keep up with open models
- musicale 2mo agoIs there any reliable way to evaluate how well "alignment" actually works?
- guluarte 2mo agoI think it's an excuse to cut R&D spending (training new models) to improve their margins ahead of the IPO. Instead they'll focus on developer growth, offering more free tier benefits, higher usage limits, etc., to expand their user base. Essentially, they're pivoting from R&D investment to profit optimization
- Havoc 2mo agoMeanwhile I can’t get a western LLM to look at a repo and tell me whether it contains anything malicious (it was a skill repo - literally just text files). Alignment my ass
- hedora 1mo agoIf you think that's bad, try asking for gardening tips.
- willrshansen 1mo agoOf course. They are slowing down intentionally because their technology is too powerful. They could totally go faster if they wanted to. No bamboozle.
- csomar 1mo agoIn my opinion, the bubble is very close to burst and they need to move quickly. I have subscribed to Claude today as I wanted to work on some amateurish CLI and then realized how much Opus 5 sucks. Surprised, I checked reddit and found that my experience is not far off from the rest. It's a massive downgrade from 4.8. I stopped using the Chinese models because even though they are workable, they are too expensive as they are not as subsidized as GPT/Claude subscriptions. Open AI use is particularly subsidized these days. A $20 sub, gives you roughly $400-500 of API use and quasi-unlimited chats. There is no way people are paying $1.000+ for a chatbot. Most people don't even pay for search. And it is expensive to run these models as the Chinese models have shown that better performance is yielded mostly from the model size. LLMs also fail spectacularly at making any decent software. I haven't seen any so far and they write fast. So we should have something by now. Crypto is getting the heads up as capital is getting re-arranged. Bitcoin/Ethereum are up 10-20% today.
- bobkingdom 1mo ago[flagged]
- beyondscaletech 1mo ago[dead]