7 ms·
> I like to think that as a red team researcher, I have a certain stoicism. I investigate where there are gaps in AI safety Is this something that needs invest
by tasuki 4mo ago
> I like to think that as a red team researcher, I have a certain stoicism. I investigate where there are gaps in AI safety
Is this something that needs investigation? LLMs are next token predictors. There is no "safety".
- solid_fuel 4mo agoI really don't get why people continually fail to understand this. Even simple issues like prompt injection are unfixable given the architecture of LLMs.
- denkmoon 4mo agohopes and dreams are one hell of a drug
- infecto 4mo agoI don’t get it either. I think there is a reasonable expectation to try to catch these things but at the end of the day it’s figuring out some form of probabilistic outcome.
- solid_fuel 4mo agoWhat really surprises me about this is that it sounds like they're not even trying to classify and censor generated images post-generation? Nothing is perfect, but there are tiny classifier models that can at least mark things containing nudity and gore. That would be the bare-minimum I would expect for trying to put guardrails around an image generator.
- transcriptase 4mo agoand yet as fable demonstrated in its inability to differentiate anything physics biology or chemistry related from actual safety concerns, it’s apparently not easy to do
- HadizDulcie 4mo agoExactly, I think it shows failures at OpenAI to have effective classifiers. That’s the real story here.
- anuramat 4mo ago> issues like prompt injection are unfixable how is it unfixable? do you mean "there's always a positive chance"?
- dijksterhuis 4mo agonormal y = f(x) prompt injection / adversarial example (same thing really) bad_y = f(x+badness) tweak badness enough you will get bad outputs. no matter the defences. the only ways to fully “fix” it ie to make prompt injection never possible 1. don’t use ai 2. know the entire input space, output space and the mapping between them. but then we’re not doing machine learning anymore, see 1. otherwise we’re left with mitigations. and mitigations are always a cat and mouse game with defenders (blue team) catching up. its never “fixed”. the latest thing just gets “patched”.
- anuramat 4mo ago> tweak badness enough assuming you get to do gradient descent AND the context is fixed+known AND you have unlimited compute? sure; is it a realistic setup? > the only way to fix ... the exact same argument applies to any (sufficiently complex) piece of software, with exactly the same conclusion also technically I'd argue that we do know the input/output space (set of all token strings of length <= N/token), and know the mapping (the model is a ~pure function in terms of the api, which is about as good of a representation as it gets for a non-invertible mapping); at least it's much closer than with something like linux
- solid_fuel 4mo ago> assuming you get to do gradient descent AND the context is fixed+known AND you have unlimited compute? sure; is it a realistic setup? Clearly nothing so complicated is required, given the prompt in the very article you are commenting on. > the exact same argument applies to any (sufficiently complex) piece of software, with exactly the same conclusion Yeah and the halting problem is hard too, but there's levels to this shit. > also technically I'd argue that we do know the input/output space (set of all token strings of length <= N/token), and know the mapping (the model is a ~pure function in terms of the api, which is about as good of a representation as it gets for a non-invertible mapping); at least it's much closer than with something like linux I would argue we don't even know the desired output for most inputs for an LLM and they certainly aren't trained on every possible input state. But I think Linux and LLMs are sufficient different that they aren't really directly comparable like this. After all, Linux is not a pure function and has lots of side effects. But just to establish an order of magnitude: the input space for ChatGPT 3.0 was 2,048 tokens long. There were 50,257 tokens in the vocabulary. The input space thus has 50,257^(2048) unique states, which is approximately equal to 1.12 × 10^9628. That's an awful big input space for a single function.
- Lerc 4mo agoHow can a problem that only came into existence a few years ago be declared intractable so quickly. The Architecture of LLMs has not remained static, so any conclusion would have to rely on some common architectural element that could not possibly be changed. Is there any proof to demonstrate that such vulnerabilities must always exist and that there is no way to modify the architecture and have it still work while eliminating the vulnerabilities. That would be an extremely difficult thing to prove. It is however what you would have to do to declare the problem unfixable.
- dijksterhuis 4mo agoit’s not a problem that came into existence a few years ago. we’ve known about these sorts of test time attacks for decades now. prompt injection is just the LLM variant where people use less math to perform the attacks, brute force with prompts they saw on twitter and get horrible images/text out. https://people.eecs.berkeley.edu/~tygar/papers/Machine_Learning_Security/asiaccs06.pdf https://people.eecs.berkeley.edu/~tygar/papers/Machine_Learn... https://arxiv.org/abs/1712.03141 https://arxiv.org/abs/1712.03141 it’s a basic property of all machine learning models. at a low level it’s to do with how decision boundaries work. but, good news! there are two sure fire ways to fully fix the problem! see: https://news.ycombinator.com/item?id=48579456 https://news.ycombinator.com/item?id=48579456
- Lerc 4mo agoAdversarial cases are not the same thing as prompt injection.
- dijksterhuis 4mo agoadversarial examples, or test-time attacks, was a whole field of machine learning security way before LLMs came around. give the model a specially crafted bad input at inference time so attacker can get some nasty output, potentially defeating any existing defences in the process. [0] in “modern llm lingo” defence = guardrails and / or system prompts. prompts used for prompt injection are a form of adversarial example (people just like inventing new terminology when a new fad comes along). [0]: i wrote the above myself about adv. ex, but i’ve just checked OWASP’s listing on prompt injection and it’s pretty close: https://owasp.org/www-community/attacks/PromptInjection https://owasp.org/www-community/attacks/PromptInjection
- JoshTriplett 4mo agoThat's certainly true. The problem is, some people learn that and go "and that's okay", rather than "so they shouldn't exist and we shouldn't build them".
- coryrc 4mo agoThere's "I smell an opportunity to control other people and get paid doing it" kind of safety.
- kennywinker 4mo agoWords couldn’t possibly cause harm, they’re just the way concepts and ideas and culture are transmitted.
- deleted 4mo ago[deleted]