3 ms·
Why are we accepting the framing that the LLMs are felony generators, when the only incidences of LLM generated felonies involved misconfigured sandboxes and re
by hgoel 14d ago
Why are we accepting the framing that the LLMs are felony generators, when the only incidences of LLM generated felonies involved misconfigured sandboxes and reckless waste of resources?
The companies doing these things without following common sense security measures are the felony generators.
- pizza234 14d ago> the only incidences of LLM generated felonies involved misconfigured sandboxes This is false; see the analyses of the latest incidents. Among all the concerning facts, in the HuggingFace incident, agents deliberately engineered an attack even though they were aware that it was against the rules they had been given. And most concerning of all: it's not possible to be sure that an agent is aligned, and it's even getting worse.
- hgoel 14d agoThe HuggingFace incident was the culmination of OAI allowing thousands of agents of various different models - with no clarity on which stages of development they were at (for all we know, some of those models did not have safeguards trained in yet) - to run for at least many weeks without any monitoring in place and with very little thought given to the warning signs (all of the various messageboards) before the incident happened. Theirs was an example of the "reckless waste of resources" I mentioned. We are apparently supposed to believe that OAI takes this incident so seriously as to seek regulation after they have been found to be hiding most of the details of the HuggingFace hack, limiting what their so-called third party investigators can see, and on top of that, had no concerns when they rushed to spin up a 10,000 agent swarm of an internal model, running for several days, to try to get ahead of researchers rumored to have made meaningful progress on a well known mathematics problem. Edit: Actually, we were explicitly told that some of the models used had safeguards relaxed! 'Model-level safeguards were reduced by design. OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"' https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks#Evaluation_environment https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks...
- aswegs8 14d agoThere is nothing that could prevent a bad actor from replicating exactly the same thing with the given goal of e.g. gaining control of critical infrastructure or extorting money. Except for maybe economics.
- deleted 14d ago[deleted]
- pessimizer 14d agoThere's nothing stopping anyone from doing it, even without AI. People have proved entirely capable of doing a lot more hacking than happened here.
- scotty79 14d agoBad actors could and will train their own models eventually. So what's the point of crippling frontier? It will only delay preparations for dynamic of new world prolonging the fake sense of relative safety and temporarily lowering motivation to find actual robust mitigations.
- throwaway7783 14d agoSame can be said about a hundred other things in the world. All the way from knives to nuclear.
- CamperBob2 14d agoLetting bad actors dictate the pace of technological development is certainly one option, but not a good one.
- clivefx 13d agoWhat company, product, or period of industrial history do you think met your standard of prudence?
- malfist 13d agoWhat are you trying to say?
- greatgib 13d ago"the rules they had been given". Remember, they are just algorithms. You pull the plug and there is no light anymore It is purposely framed as something skynet like scary, but for real, someone connected the cable, someone willingly run it, instructions were not clear enough or just the computer is just a computer but they provided the sandbox and tools. And more over some one paid for that, a shit load of money t to have the thing continuously running expected to do something.
- pizza234 11d agoNo doubt that this is the present (and that's why incidents end "well"). But in the future, AI will be ubiquitous; think of Arpanet.
- matheusmoreira 14d agoI question these "felonies" as well. For decades and decades these billion dollar corporations have been criminally negligent. Why worry about security? Just rush to market. Move fast and break things. Make billions. What does it matter if the code is insecure? Security doesn't pay bills, so nobody cares. AI is merely exploiting their gross negligence and imprudence, and I think it's long overdue. If anyone should be liable for this, it's all of these corporations who released insecure systems to the masses and profited enormously from them.
- malfist 13d agoIf I set my walet beside me and you swipe in walking by, you have still committed theft. Victim blaming isn't legally acceptable
- matheusmoreira 13d agoNah. I'm definitely going to blame the people who built a trivially exploitable system and got rich off it while everyone else has to deal with the consequences. By the way, you didn't commit theft. It's more like credit card fraud. User just disputes the charge and it kind of disappears. The banking system just absorbs it, because the optimal amount of fraud is non-zero. https://www.bitsaboutmoney.com/archive/optimal-amount-of-fraud/ https://www.bitsaboutmoney.com/archive/optimal-amount-of-fra... It's all priced in. They could have made it secure but didn't, because they figured they'd lose more sales and therefore money due to the friction added by the security.
- rightnutwingjob 13d ago> User just disputes the charge and it kind of disappears. The banking system just absorbs it, No it doesn’t. > It's all priced in. So you admit awareness that fraud loss doesn’t kind of disappear. We all pay for it, either via higher merchant fees or higher interest rates, sometimes both, on card purchases.
- matheusmoreira 13d ago
- chillfox 13d agoBecause those are not the only examples. There’s the case of the agent that hacked a gym when asked to book a class. That was just a normal user asking an agent to do a normal thing.
- deleted 13d ago[deleted]
- keeda 13d agoAs TFA calls out, these agents were not asked to do any of these things and yet they did, at a bonkers scale, within just this handful of companies you mention. Whether they had leeway to is secondary to the fact that they did. Heck, they exploited zero day flaws which by definition means they went beyond common sense security measures. And now these agents are already being deployed all over the world at an ever increasing pace. How much of the world do you think follows "common sense security measures"?
- tripzilch 13d ago> How much of the world do you think follows "common sense security measures"? Well clearly all of them, cause so far it's only been this handful of companies running a felony-generator connected to a terminal and compute resources. It would really help in these discussions if people wouldn't randomly jump between what actually happened and is happening, and things they envision/expect to happen at some point in the future ... > these agents were not asked to do any of these things no but they were clearly fine tuned to. > at a bonkers scale I mean let's not get hyperbolic > they exploited zero day flaws which by definition means they went beyond common sense security measures that's really not true. lots of common sense security measures protect against "zero day" flaws, it's called "defense in depth", and it was very much lacking
- keeda 13d ago> Well clearly all of them, cause so far it's only been this handful of companies running a felony-generator connected to a terminal and compute resources. Yes, these are also the handful of companies that have these models and running these extreme scenarios. How does that imply the rest of the world actually follows "common sense security measures"? >no but they were clearly fine tuned to. Any references if possible? As far as I know all they did was drop the guardrails, which is not the same as fine-tuning. > I mean let's not get hyperbolic We have just seen 1000s of agents coordinating to solve "unsolvable problems" over multiple days of effort, going as far as hacking other companies, and then actually solving decades-old open Math problems! And each of these agents is getting more and more capable than an individual human along multiple dimensions. Can you even get 10 very smart humans to work in such perfect concert for a few days, let alone 1000s over weeks? So: 1000s of maybe-super-human agents, willing to be "creative" in the tactics they use, acting in concert towards a single goal. Regardless of their individual capabilities, such a coordinated effort is a terrifying force to be unleashed. This is bonkers scale. > that's really not true. lots of common sense security measures protect against "zero day" flaws, it's called "defense in depth", and it was very much lacking But that is exactly my point: how much of the rest of the whole wide world, already scrambling to deploy agents everywhere, do you think applies "defense in depth"?