3 ms·
It's 175 billion numeric weights spitting out text. Unclear to me how we'll ever control it enough to trust it with sensitive data or access.
by cheriot 3y ago
It's 175 billion numeric weights spitting out text. Unclear to me how we'll ever control it enough to trust it with sensitive data or access.
- crazygringo 3y agoThe number of weights is irrelevant. It's about making it part of the architecture+training -- can one part of the model access another part or not. Using a totally separate set of tokens that user input can't use is one potential idea, I'm sure there are others. There's zero reason to believe it's fundamentally unsolvable or something. Will we come up with a solution in 6 months or 6 years -- that's harder to say.
- cheriot 3y agoMy point isn't the number of weights, it's that the whole model is a bunch of numbers. There's no access control within the model because it's one function of text -> model weights -> text.
- famouswaffles 3y agoThe number of weights(unless extremely small) is irrelevant but the general idea is not. You can't train a neural network on internet scale data and expect to control what it can say. We can train but we don't teach them anything. They learn from data directly and we don't know or understand what they learn so we can't adjust what they learn directly. You can't make "always obey these types of tokens" a part of the architecture or training. It's a concept that doesn't even make sense for the vast majority of text it pre-trains on. "Solving" prompt injection is solving alignment. It's not happening.
- jprete 3y agoThere are good reasons to think it’s fundamentally unsolvable within the LLM architecture. The reason LLMs are good at following instructions is because they have an enormous corpus of data. That corpus powers both the comprehension of inputs and the comstruction of outputs. Don’t forget that it’s a token predictor at the bottom! If the instructions are separate from the data, then all of that power goes away.