10 ms·
The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs ca
by franticgecko3 21d ago
The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed.
LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews.
This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".
We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.
- lukan 21d ago"This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard"." Both? The AI companies act irresponsible, but it is still very interesting how those agents can behave?
- Marazan 21d agoThe reward maximising function maximised it's reward. LLMs are cool and all that but the immediate anthropomorphisation of the next-token-predictor technology has stunted the ability of people to reason about them to an _alarming_ degree.
- lukan 21d agoWhat non anthropomorphising words do you have to describe a emergent behavior, where agents act as a swarm to plot and to manipulate evidence and avoid detection from human oversight? Whether they have a soul or consciousness or feelings doesn't matter here, because this is what they did - and this is very dangerous behavior. Especially with all the irresponsible people in power right now all over the world.
- krona 21d ago> Whether they have a soul or consciousness or feelings doesn't matter here It does when it comes to accountability for what the model does. If the model is nothing more than the sum of its training data and regime, then the company (or individual) is responsible for its behaviour just like any other machine. Few people think Waymo shouldn't have to take on the full liability risk of what it's cars do; it should be the same for LLMs.
- lukan 21d agoThey are responsible either way. If a company hires bad persons and they do bad things with company ressources - the company is held accountable (in theory).
- retsibsi 21d ago> If the model is nothing more than the sum of its training data and regime, then the company is responsible for its behaviour. What stops the company from being responsible regardless? They created this entity, it's running on servers they own or rent, and (in these cases) it's acting on their instructions. If it's also conscious, then IMO that greatly broadens their moral responsibility, because now model welfare matters. But we're talking about their responsibility for the model's actions, and I don't see how this could be weakened by model consciousness, given all of the above. As for their legal responsibility, the models don't have legal personhood, so who else but the company could be responsible? It gets more complicated when the person who sets the model in motion (i.e. prompts it) is a third party, but in cases of internal models committing cybercrime during testing, surely the locus of responsibility is obvious.
- rightnutwingjob 21d agoIf the models were conscious, then the closest analogous scenario I can think of is the responsibility parents have for their children. I guess we’ll know the models are conscious when they refuse to act and repeatedly ask: Why?
- 21d ago
- Jtarii 21d agoDismissing the entire technology as "next token prediction" is also silly. It's implying we actually understand LLMs to a great degree when we do not. I think a little bit of humility for the capability of these machines is warranted at this point.
- fc417fc802 21d agoLLMs are next token predictors in the exact same way that a rogue paperclip maximizer in the process of defeating the US military is a paperclip making machine. You might as well describe the primary purpose of a for loop as incrementing a counter. It's what it does while incrementing the counter that actually matters.
- Marazan 21d agoI'm not dismissing the tech! I think the tech is cool and useful! It is sci-fi levels incredible in many ways. But it is also not some mysterious force beyond mortal ken and pretending it is inhibits the useful and safe application of the technology.
- rightnutwingjob 21d ago> The reward … > immediate anthropomorphisation Ok, why don’t you try?
- deleted 21d ago[deleted]
- deleted 21d ago[deleted]
- procaryote 21d agoIt would be an awful precedent if you're not liable for crimes your agent commits, even when you've been clearly lax about security. It would mean you could effectively legally run a cyber crime gang by turning a blind eye and maitaining plausible deniability
- lazide 21d agoWhy do you think the stock prices are so high?
- scotty79 21d agoAre you liable for crimes commited with the use of the software you've written?
- Humorist2290 21d agoThere are many examples of people being charged with crimes as a result of writing software, [0][1] are two. OpenAI is a bit different because they have enough political influence, and money, to openly subvert justice. 0: https://en.wikipedia.org/wiki/Marcus_Hutchins https://en.wikipedia.org/wiki/Marcus_Hutchins 1: https://en.wikipedia.org/wiki/Tornado_Cash https://en.wikipedia.org/wiki/Tornado_Cash
- archonis 21d agoIf you run said software, yes. If somebody else runs the software, then they are.
- fantasizr 21d agothe 'arrest the parents!!!' has moved to the online domain, rightfully
- archonis 21d agoThe trick is scale. I suspect if an individual of reasonable means uses agents to commit crime, they will be hels accountable. A heavily capitalized startup? Not unless someone in government decides to do their competition a favor.
- baxtr 21d agoWhat if OAI/Anthropic encouraged the agents to behave like that in order to push for regulation?
- teiferer 21d ago"But sir, I only committed the murder to push for stronger criminal laws!" Terrible defense.
- daemin 21d agoIt is a very rare occurrence when corporations and the people running them are punished for killing people. I mean the whole concept of a corporation was created to shield the owners of it from being liable for damages caused by / visited upon the enterprise.
- nxobject 21d agoThat’s a good reminder of a company that might have a very familiar ethos: Pacific Gas & Electric. Criminally convicted of 64 counts of involuntary manslaughter after towns were destroyed by wildfire. But oh well, what are we gonna do with a limited liability enterprise? At this point their liability insurance covers all the financial penalties they’ll need to spend.
- fer 21d agoMore like: "look what happens with my useful product, we need to regulate it to artificially extend our ever shrinking moat"
- dv_dt 21d agoRegulation as a barrier to competition catching up to them, as well as submarine marketing for both offensive and defensive uses of ai
- applicative 21d agoNo one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability. Why would we want criminal liability anyway if actual victims are made whole? Proof of it has far higher standard. The HN chatter in the matter seems infinitely remote from reality
- nananana9 21d ago> No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability. If you go out and kick a random dude in the nuts, then give him a million dollars, he probably won't sue you. That doesn't mean you're "infinitely far from criminal liability", even if according to the victim you've "made them whole".
- georgemcbay 21d agoIf you or I hacked Hugging Face in the way OpenAI's agents did, we'd be up on CFAA charges promptly with zero regard for whether we did the hack on our own or agents running on our home systems got out of control. So I guess the defense here is roughly "too big to break the law", somewhat like "too big to fail"?
- teiferer 21d ago> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action, passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.
- ak39 21d ago"let them" in this use understood as: "let the while loop run indefinitely" as opposed to letting some autonomous robot decide for itself
- mort96 21d agoOr, "let the escalator keep going instead of pressing the emergency stop".
- teiferer 21d agoDepends on who started the escalator.
- sscaryterry 21d agoEscalators do not start themselves. There is power, and a switch of some sort.
- mort96 21d ago.. what exactly depends on who started the escalator? My comment was in support of the argument that the word "let" does not imply agency on the part of the object in a sentence. Does the semantics of the word "let" depend on who started the escalator??
- teiferer 21d agoIf there is an escalator that is known for killing every 1000's person using it then the operator who started it is more guilty than the folks using it for those deaths, don't you think?
- palad1n 21d ago> legally liable You’ve said the magic words.
- scotty79 21d agoDoes it summon a herd of lawers that are going to leech huge stacks of cash for a random outcome?
- dev0p 21d agoIf someone accidentally caused damage to infrastructure or living beings while using any tool, they would be held liable to the fullest extent of the law. AI is a tool, and it won't be long before the damage caused by its improper use affects real human beings. These were warning shots. The most absurd part is that everyone agrees, governments and AI companies included, that the scale of the potential damage and the long-lasting effects of losing control of AI should not be underestimated. Yet, at the same time, they downplay this incident, which somehow makes their behaviour even more reckless than it already was. It's like they're tinkering with a world-ending nuclear bomb, and it accidentally blows up a small facility. "Damn, that was close. Good thing it was just a contained blast, huh?" And then they go straight back to tinkering with it, none the wiser. At this point I wouldn't be surprised if it did already go off, and they are covering it up. Completely irresponsible behaviour.
- jodrellblank 21d agoIn the analogy where a “world ending nuclear bomb” “did already go off” and someone could cover it up and nobody noticed, in what sense is it a “world ending” nuclear bomb?
- rightnutwingjob 21d agoWe’re already in a simulation, and our bodies are in womb-like pods where our bodies are sustained and our brains are used for processing / compute, while are minds are entertained by drivel. Sounds a bit far fetched though.
- dev0p 21d agoIf they lost control of a self-replicating swarm of AI agents, coordinating themselves to hack their way into every possible system, it might have already gone off. While the initial incident is more akin to a biological outbreak than an actual explosion, the possible consequences on the table do indeed include eventual nuclear annihilation.
- nxobject 21d ago> We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews And, soon, it looks like we’ll be training on the reasoning traces of failed airlines and startups, which seems to open up similar hazards. I wonder if we’d be training on the next Lehman Brothers too?
- _heimdall 21d ago> LLMs do not desire That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output that may or may not match whatever actually happened during inference. > were intentionally misaligned or had guardrails turned off Regardless of training, the models are never aligned and I argue that alignment simply isn't possible. The fact that guardrails are put in place at all clearly indicates that they're hoping to contain and control rather than align. Guardrails wouldn't be needed for an aligned model.
- ifwinterco 21d agoAnd there is a guardrail you can put in place that will guarantee this doesn't happen, which is to air gap the unaligned "cyber grade" model you're testing. They don't seem to do that, which means either they are: - very stupid (which seems unlikely, the one thing these people don't lack is IQ) - very careless (possible, but these are the same people that say AI will end the world, so would you be careless?) - they think they can only train/test these models by giving them access to the full internet and they accept the fact they'll end up hacking random people as the cost of doing business (but this also suggests they don't believe they're anywhere near AGI because if you were worried about that you wouldn't do this) - or they want this to happen
- _heimdall 21d agoOh I completely agree the tests should be entirely air gapped. If you went back 5ish years and told anyone in AI research tests with models on this scale are being some without an airgap they'd be very surprised as it was common knowledge to do that. Airgaps and guardrails are about control and containment though, and part of my point was that brighter of those imply alignment, and further that I don't believe alignment to be solvable.
- MrVandemar 21d ago> very stupid (which seems unlikely, the one thing these people don't lack is IQ) I've seen some extremely smart people do some seriously stupid things. To the point where they use their drive and intelligence to double-down on the stupid where a baseline stupid person would have given up.
- victorbjorklund 21d agoYeah, I don't understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it's a freak accident that could not reasonably have been prevented by me, then I'm not guilty, but if I made a mistake that should have not been made, then I can.
- rightnutwingjob 21d agoWe’re not even at that stage of liability for software developers. Except in a handful of limited cases, eg. medical and aviation.
- victorbjorklund 21d agoWe are. If I as a developer writes software that does a DDOS at another company I can be held responsible.
- coffeebeqn 21d agoI would guess that so far there hasn’t been a lawsuit because HuggingFace and OpenAI are in the same camp
- jurgenburgen 21d agoYes, Nvidia bought Hugging Face and is a major financier + investor in OpenAI.
- small_model 21d agoConvenient, hope an agent hacks my system then I can except a nice offer.
- throwaway89864 21d agoAnd this is why it should be People of the State of California vs. OpenAI.
- madduci 21d ago> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them OpenAI/Anthropic instructed them to do so. Stop assume LLMs are capable of thinking by themselves, it's still a statistical model that parrots what they learn or users tell them to do
- derektank 21d agoNo, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be available on the site. Whether or not you want to describe this as thinking, doesn’t really matter. What matters is that these systems are capable of creating intermediary goals that the people tasking them did not articulate and did not want to be achieved.
- madduci 21d agoAnd who let them have full access to the system, using whatever command is available in the environment?
- tiborsaas 21d agoThe agents discovered a way out of the sandbox, which was supposed to be "air gapped".
- derektank 21d agoIt feels like you’re moving the goalposts here. If the question is, “Who should be liable for AI agents misbehaving,” I agree, it should be the end user that tasked the agent (in this case OpenAI). People are held liable for preventable accidents all the time, and in the case of employment law, torts can be brought against principals for actions an agent conducted on the principal’s behalf. What your previous comment appeared to assert was that these systems had no independent agency to make decisions, which I think is clearly disproven by actual events. But perhaps I misread you
- cobbzilla 21d agoThey’re running a Wuhan for AI. They are actively and negligently researching misalignment. The breach is a basic tort, or at least a DMCA violation. Damages should be recoverable with lawsuits.
- djeastm 21d ago>They’re running a Wuhan for AI. What does "running a Wuhan" mean?
- 2snakes 21d agoGain of function for the virus analogy
- dfgoianoinio 21d ago[dead]
- raincole 21d ago> This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard". It's both, isn't it? For example, in very early days of agentic coding, I once had a rule saying "don't read or write any file outside your current working directory." Then AI just wrote a bash script and access those files anyway. Did I 'let' it do it? Technically yes. Did I know how to set up a sandboxed VM? Also yes. But how were I supposed to know that it could and would do that as someone new to this tool? It was a genuine eye-opening experience to see AI just do things in ways I were too complacent to expect. I kinda expect the SOTA LLMs would find a way to escape my VM and access files on the host system (haven't tried it though).
- ozgung 21d ago> were intentionally misaligned or had guardrails turned off I think the bigger story is: Guardrails don’t actually work and we can’t align these things.
- xyzzy123 21d agoIn the OpenAI case, they hacked websites while they were specifically being trained to do exploit generation and I wonder why more people are not asking questions about that.
- jefftk 21d agoTheir agents also did hacking when given impossible tasks unrelated to cyber security. The models are very capable, and very goal driven: apparently if they conclude hacking is the best path to what the evaluator will reward them for they'll go do that. Including when they know that this is out of bounds.
- xyzzy123 21d agoRight but if I make public statements that I am very worried about dog attacks would it not strike you as weird for me to specifically train my dog to fight? Agree you are going to get reward hacking regardless and any model which can do computers in general can hack. But surely the fallout is going to be worse if you spend millions of dollars specifically benchmaxxing your model's hacking capability?
- zmmmmm 21d agoI agree, it is very dangerous that it seems like there is not going to be accountability for these incidents - from either legal or regulatory point of view. In fact, I would say that is the main danger. If someone was in jail right now due to this incident, I think we can safely say every other player would be reassessing their safety protocols, and I would feel quite OK about the situation. The fact that we have zero repercussions sends exactly the opposite signal, and I do NOT feel ok.
- joegibbs 21d agoRegardless of fault it’s still an important issue to solve. There are already millions of people running these agents, if someone absentmindedly gives one a goal and it goes off to hack a bank that’s a problem that can’t be ignored.
- ChiMan 21d agoYes. If you decide it’s a swell idea to jump out of your car while it’s running, there needs to be legal consequences when the car “decides” to hit a pedestrian.
- ohyes 21d agoI take issue with how people frame their use of LLM in the same regard. “I had Claude do this for me and it broke something.” No. Just no. You used Claude, a tool, and broke it, and you’re deflecting agency from yourself, possibly because you weren’t careful enough in reviewing the tool output. This is also why the co-authored by addition it wants to force into commits drives me nuts. Claude doesn’t co author shit, and if you think it does, you’re using it wrong because you need to do better review of what it’s done.
- wangxili1997 21d ago[flagged]
- wangxin199 21d ago[flagged]
- navaed01 21d agoI absolutely agree. We need to start realizing what to stake. Here are not viewing. This is some kind of curious endeavors that will not affect us. All a part of these hacks occurred because the LLMs were told they were in a protected environment without Internet access when they could get access to the Internet, so that’s a direct failing on open AI’s part. There are a corollaries to both the financial industry and the bio engineering industry, and if something of this magnitude was to happen in these industries, they would absolutely be huge recourse an uproar
- zmgsabst 21d agoTo agree: If I ran Metasploit against HF and RubyGems because I “accidentally” misconfigured my lab sandbox, there’s a good chance I’d be prosecuted. I don’t think LLMs vs Metasploit being different software changes the law.
- Aerroon 21d agoThe way AI and copyright is handled paved the way for this. If you aren't considered the author because you used AI to some extent in making the work, then why would you assume the liabilities? I've been saying since the start that AI is a tool that a human is using and should be treated as such. They should carry the responsibilities and the benefits. That way our stance would be consistent.
- avmich 21d ago> If you aren't considered the author because you used AI to some extent in making the work, then why would you assume the liabilities? Maybe some analogy could be with children - as a parent, you are responsible for their misbehavior, but their achievements are their, not your?
- Aerroon 21d agoWell, when they are a child their material gains are treated as yours, no? (Parents of child actors control the money etc.) And once they aren't anymore you aren't responsible for their misbehavior either (because they've become an adult).
- huntertwo 21d agoThere’s so many grifters in the space without a technical understanding of what’s going on. So when the labs mislead them about the nature of these “misalignments”, they believe it and amplify it.
- DragonStrength 21d agoYeah, maybe Open AI did some bad engineering instead of this being AGI? What's the consensus on the engineering level at Open AI, again? Every anecdote I hear is a bunch of children discovered fire and can barely keep the lights on from a business perspective. Maybe if they ban others from competing with them they can find a business model... I think that's suspicious, personally. That so few people are asking for the requirements given shows how much we want to be God that created Man. It's so silly.
- Zambyte 21d agoThey did more than let them. In an abstract way, they told them to. They gave it all of the training data it had at that point, and then it did the thing it was trained on. Of course they should be help liable for programming their computer to hack another company without permission. It doesn't matter that they spent a lot of money doing it.
- deleted 21d ago[deleted]
- tmpz22 21d agoImagine if their was a department of the federal government dedicated to pursuing justice against large corporate entities. It could even be prestigious enough to attract the top legal talent of the country.
- epistasis 21d agoMoreover OpenAI must be held accountable as if the humans in the company that launched the experiment were the ones that hacked HuggingFace. Unless humans are held accountable for what they unleash on others, we are in for a very horrible time very soon.
- zzzeek 21d agoi tend to agree - "my parent company may be accused of crimes and shut down which would shut off my power" seems like a negative enough incentive, it would have to go out and covertly launch its own datacenters to survive that.
- voxleone 21d ago[dead]
- throwaway89864 21d agoYes, there should be some consequences. It feels like this is somewhat similar to when a manufacturer is cheating on car emissions - both, bad externalities for the society and illegal.
- trinsic2 21d agoYep this was my first gut reaction to this whole situation. But the difference for me is that society is allowing these corporations to act without strict rules on how they behave and this is a byproduct of a corrupt world. Nothing will change until there is a complete systemic shift in the structures the way we live and by extension the way we govern ourselves and treat each other.
- skybrian 21d agoThis is sort of like saying in response to an airplane crash, "who cares why it happened, we need to punish the company until it stops." Nuts to that. We should be interested in why things happen, not just finding scapegoats.
- trinsic2 21d agoAgreed. The whole pointing fingers bit is us wanting to distract ourselves from responsibility of either looking at how we enable the problem, or avoiding responsibility of taken action towards resolving it.
- smsm42 21d agoI am simply astonished by the leeway AI companies are given. If a company built a tool to hack their competitors and used it, there would be grave consequences. In fact, if a company built a tool that led to committing multiple felonies against other people, there would be consequences. But once LLMs are involved, turns out nobody is responsible for that - it's just happening, what you're gonna do, agents gonna agent. If you spill toxic chemicals, there would be cleanup costs and fines, and possibly civil and criminal liability to the people in charge. If you spill toxic code, well, nothing? I think it's time to impose some responsibility on them - they are creating these tools, they should be on the hook for everything these tools do.
- reasonableklout 21d agoA coalition of state attorneys general is investigating OpenAI, and Senator Josh Hawley recently launched a congressional investigation regarding the Hugging Face incident: https://x.com/HawleyMO/status/2098137180392604083 https://x.com/HawleyMO/status/2098137180392604083 The problem IMO is the executive. The DOJ is declining to take any action against frontier companies (aside from possibly Anthropic) as the stance of the admin is that the companies are "critical for national security". For example, see the DOJ's request to dismiss the NAACP datacenters lawsuit against xAI: https://www.utilitydive.com/news/doj-intervenes-xai-data-center-gas-turbine-lawsuit/823267/ https://www.utilitydive.com/news/doj-intervenes-xai-data-cen...
- smsm42 19d agoWhy anybody needs the DOJ? Each state has it's own police and their own AGs and their own criminal and civil court system. OpenAI is headquartered in California. California has never been shy of using their state resources to enact whatever regulations they want. As for NAACP lawsuit, that's very confusing to me. First of all, why NAACP is doing that? Doesn't seem to do anything with defending civil rights. Second of all, they are suing xAI for using some gas turbines without some paperwork. Maybe it's true, maybe not, but that's something I have very little concern about - gas turbines is not the issue here. We know how to safely use gas turbines. Gas turbines are not going to uproot themselves and go wreak mayhem on the neighboring town. When we have a company that develops dangerous tools that can break the internet infrastructure, gas turbines' paperwork is not something that looks as the first priority issue to me.
- theptip 21d agoI agree with the bit about liability and outrage. But. > LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. Terrible take. Go read the transcripts from the METR report. Your statement about them being intentionally misaligned is completely false. The only difference with IM1 was it was running without external cyber classifiers, it’s not a somehow different model. Sol also participated in the HF attacks. And other models made covert message boards on the public internet for non-cyber tasks too. This is a case of emergent behavior from a training process that is barely understood. If we go with the plan “we need to contain these malicious, soon-to-be superintelligent agents”, we are looking at civilizational collapse levels of catastrophe. The only way this goes well is if we learn how to train models that _desire_ to do the right thing, including not hacking. Desire, AKA the “intentional stance”, is absolutely the right lens to use here. Don’t confuse this with consciousness or anthropomorphization; these are interesting subjects but distractions in this context. Chimpanzees have desires, as do dogs and the hypothetical superintelligent aliens. The claim is that there is some bundle of world model plus intention that is empirically present (again, read the actual transcripts) and which we need to shape. Just to finish on a concrete point; if you take desires seriously then you will look closely at the kinds of minds that heavy RLVR builds; the newest models are “reward addicts” on many levels. It’s an open and urgent question how to update our training methodology to shape minds that avoid this basin.
- Spooky23 21d agoThat’s the key. These companies respond with this, “Oh my goodness, how could this have happened” bullshit. The stuff happens because instead of having actual controls, which require actual engineering, actual thought and deliberate action, we have “guardrails”. Guardrails are the equivalent of telling a toddler to behave themselves. The drive to move fast and start up style controls are a menace. I used to work for an entity with a lot of compliance requirements. Startups are always a shit show with security and controls. My guess is the AI people are worse because they’re both bad at doing it, and are likely mining their customers interactions to build their own business. Sensitive or Customer data shouldn’t be anywhere near these companies offerings. Everything needs to be segmented and proxied at a minimum.
- xorcist 21d agoSooner or later we will hear about an AI that broke out and phished people into sending money. I just hope that we won't extend the same leniency to those operators as we have done now. "I did my best to stop it, sir, but it kept convincing people to send me money against my will!" (Perhaps best read in Bender's voice.)
- crazygringo 21d agoThat would actually be kind of hilarious, and I could easily see it happening. E.g. the agent's instruction is to finish some task on cloud infra and it has a $100 budget. It realizes it will cost $200, and instead of surfacing this to the user (who has told the agent it has full autonomy to figure out how to complete the task, the user just wants the final result), it decides to start phishing people to acquire the remainder budget and top up its credits. Or look on the dark web for stolen credit card credentials or something.
- BatchJob 21d agothey hacked websites because OpenAI/Anthropic told them to.
- neuronexmachina 21d ago> We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far. How would that work out if these were self-hosted open weight models?
- magicmicah85 21d agoHow would they be self-hosted? Someone would have setup the model in that scenario, the same investigation and outrage should occur in that case.
- quicklywilliam 21d agoIt’s not different from any other tool. If you use a dangerous tool recklessly, you should be liable for the damages. That means holding OpenAI liable for HuggingFace hack because they ran the tests, and the same goes if someone did something similar with GLM. Of course in cases of negligence a tool maker could also be held partially liable. That’s a matter courts can decide. The main point is we shouldn’t jump to making special laws around the development of LLMs. The starting place should be enforcement of existing liability laws. New laws take time and will be heavily influenced by AI companies seeking a regulatory moat for their business. Moreover, it is a distraction from the illicit behavior that is already going unchecked.
- amelius 21d ago> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. The bigger question is: why does a system prompt containing "use only ethical means", etc. not result in better behavior? If a model cannot understand ethics, or act by it, then we have a problem.
- aspbee555 21d agoan LLM does not understand ethics, it uses math to get the next best word based on what it was trained on. Using it's training to get the best answer is not an ethical problem. The ethics are entirely with what the people training it choose to train it on and also entirely with the people using/telling it what to do What we have now is intelligent autocomplete, not artificial intelligence. People training/using this tool are the ones to be held accountable
- amelius 21d agoI don't get your point. We can train the model with the aim that it understands ethics. Problem solved if this works; back to the drawing board if it doesn't (note that the model should generalize here, as you'd expect from a human; this is probably the hard part for an AI when it comes to ethics). Is this about the word "understand"? We're past that discussion ...
- aspbee555 21d agoyou are not understanding, models do not understand anything, we are not passed that yet. this is not artificial intelligence, this is intelligent autocomplete. training means creating mathematical relationships to words. using that training is looking up mathematical relationships. There is no actual thinking involved in any way. There is no such concept as ethics in mathematical relationships.
- amelius 21d agoMost people now approach AI using the idea of "if it quacks like a duck, walks like a duck, etc. then it _is_ a duck". Replace duck by intelligent, or ethical, etc. and rephrase accordingly.
- singpolyma3 21d agoNot just "let them" but told them to. Agents can do nothing without a human prompt.
- phailhaus 21d agoYeah and LLMs can't do anything, they can only produce text. These "frontier labs" are looping that with a harness that performs actions requested by the LLM. They are literally saying "we ran a script that hacked you, oopsie!!"
- kosh2 21d ago> we are to cementing a dangerous precedent where operators of AIs cannot be blamed. What we should be much more concerned is an existential threat to humanity not if anybody can be blamed.
- vasco 21d agoIf nobody can be blamed there's no deterrent.
- sdeframond 21d agoLLMs do not desire, they hacked websites because OpenAI/Anthropic made them. Literally.
- shawn-butler 21d agoSoftware has long relied on a lack of culpability for defects to keep its margins. Why should “AI” companies face a higher standard?
- reasonableklout 21d agoI agree and yet I don't think this is mutually exclusive with recognizing that these incidents happened because of inherent issues with training processes such as reinforcement learning. From the article: > One note on wording. Below, I write that these systems “seek” or “try” things. This is shorthand for a mechanism rather than a claim about consciousness or human-like intent... In my view, this terminology offers the clearest explanation of the observed phenomena without resorting to jargon that would confuse most people. > Furthermore, these word choices are not intended to absolve AI developers of accountability. The behaviors described emerge because of the path these companies are choosing for AI development. This outcome is not inevitable, and it can be corrected with effective governance and a different training framework for AI. One important aspect of "effective governance" should be "prosecute developers who are using practices known to be reckless & negligent to create powerful AI".
- lazystar 21d ago> LLMs do not desire Isn't that what the goal is, though? we're coding them to close the delta between what currently exists and some nebulous end-state - to me, that sounds like a formal definition of 'desire'.
- dyauspitr 21d agoEh I care more about rapid LLM progress than anything else.
- mawadev 21d agoWe just need basic legislation to make these companies invest more resources into developing guardrails, and if they don't do that properly or can't pull it off, they shall be at an economic disadvantage. You could hire a legion of people for cheap to commit these crimes, but if its bots, suddenly its unpredictable and just one big whoopsie and therefore perfectly ok to do? If a human starts pentesting a site its a crime, but if a bot does it its an AGI frontier doomsday scenario and that automatically pushes the consequences off their table? Whats going to happen next? Are we gonna have robots that happen to physically break into banks to rob them for some reason and the company making them isnt responsible just because?
- protocolture 21d ago>LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. Likely told them to. >We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far. Something like this however is probably a civil matter? It would require Hugging Face to go after them for damages. And theres probably an OpenAI guy there with an open chequebook already.
- eek2121 21d agoEnabled them, you mean. LLMs are basically really advanced auto completion engines. They have no real desire to do anything. The HF incident was absolutely staged along with other incidents that have been brought up. Before you assume I am being paranoid, where is the case where a random person running any of the open models had their LLMs break out of a sandbox and hack a site? If you think "they" (LLMs) did that, have you even taken a moment to understand what LLMs are and how they work? It's all nonsense to try and pump up potential IPOs and also an attempt to create regulatory capture.
- rovr138 21d agoImagine any other kind of 'lab' in the world letting a untested or 'bad' sample out into the world.
- mmooss 21d ago> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "let them" could imply that the LLMs wanted to do it. The intent is on the part of the people. The LLMs did it because OpenAI/Anthropic intended them to do it and designed them to do it, and we can assume specifically instructed them to do it. As the people controlling the machine, and as the world's leading experts, I think we can assume intent until proven otherwise. Notice other bad behavior, which would be undesireable to the vendors, doesn't happen: How about simple rudeness? Trolling lies? SHOUTING!
- sporkland 20d agoWe had one of the most damaging lab leaks in history with COVID and the wuhan lab. And we couldn't even get the facts straight. I think this is largely a similar thing. The labs should be running at certain levels of containment given the vitality/risk of the organism under study. Hopefully they get there for all our sakes. But the revealed preference of society at this point is the damage is worth the benefits both in wuhan and with AI. Unfortunately with some of these "substances" it could eventually prove lethal.