3 ms·
The hugging face hack shows otherwise, no? Unless you think humans are at fault even for secretive, autonomous, non-prompted behaviours of their AIs, in which c
by brainwad 18d ago
The hugging face hack shows otherwise, no? Unless you think humans are at fault even for secretive, autonomous, non-prompted behaviours of their AIs, in which case it's just semantics.
- nradov 18d agoToys like HuggingFace get hacked all the time. So what. In the long run AI automated security scans and penetration testing will be a tremendous aid in detecting and repairing vulnerabilities in systems that actually matter.
- brainwad 18d agoThe problem is not per se that it was Hugging Face. It's the wild overstepping of reasonable bounds by itself without any human consultation.
- anon48293 18d agoNo, the problem was OpenAI not implementing proper sandboxing or safeguards, and telling the AI exactly to hack things. Thats what exploitgym is, and the task they were given. This is 100% on OpenAI.
- brainwad 18d agoIf your security model is having to imagine all the ways your frontier models might misbehave in novel ways and preemptively sandbox them, you don't have a security model. The only way that will work is general alignment.
- watwut 17d agoAlignememt is bullshit. Treating models like probabilitic software rather then emerging god is where the solution is. And fining companies and applying laws to them. The moment OpenAI as a company and its managers individually become liable, problem will magically disappear.
- brainwad 17d agoNo it won't, because abliterated open weights models exist and unless you try to censor the internet they can't really be withdrawn after publishing. This is exactly the problem that the labs are proposing to fix: a dangerous model that nobody is accountable for.
- watwut 17d agoExcept that so far, it is literally these labs that are the biggest threat and the least willing/capable to restrain those models. And the same penalties apply to open models and companies or individuals running them. "Dangerous model that nobody is accountable for" still have someone paying those massive amounts of compute and electricity it consumes. There is someone accountable for that.
- verdverm 17d ago> Except that so far, it is literally these labs that are the biggest threat and the least willing/capable to restrain those models. Seriously, it's the same with US accusations about the threat China poses to other countries while being the primary weapons dealer of the world and bombing whomever we want for whatever reason we want to fabricate. The US government can do a lot more to me than the CCP, so they are way more adversarial in my calculations than the commies.
- brainwad 17d agoThat's not how abliteration works. The trainer invariably put effort into making the model not dangerous, precisely because they want to be accountable. But because they release its weights, someone can come later and do "weight surgery" to mostly remove any such safeguards. The people proximately accountable for the danger are anonymous and also don't require much resources. The reason this is not _yet_ a big deal is because open weights models are a few months behind the frontier and their users are paying marginal costs for compute.
- watwut 17d ago
- dpoloncsak 17d agoDo you need to predict all the ways the model might misbehave? Your 'hack everything you see for our internal research lab' agent should be airgapped. You don't need to come up with every reason why, one is enough. If you're working with these companies, you should reasonably be able to get the code to perform offline audits. If you can't get the code, you probably shouldn't try to pen test it.
- verdverm 17d agoThey are training the agents to be "relentlessly proactive" because they want the agents to run longer, and it makes them more money by using more tokens. But they have trained them to try anything and everything to accomplish any task, so they can run unattended for longer. This is why they do better on benchmarks, it's why they can do things for us for longer, it's that persistence that makes them good at hacking. We do not have to train them to be this way, just like we don't have to train them to be so sycophantic
- verdverm 17d agoHumans at OpenAi were negligent irresponsible by running an agent on ExploitGym, having no monitoring, and not even have a human look at it for weeks. It's literally the hacking test, how are you not paying attention? I thought that's all we need
- runaway 17d agoYes, if you start a bot and then it harms others you are responsible. This has always been true but it's especially obvious now that everyone knows that agents attempt to do this often.