3 ms·
Jesus, this is a silly take. Lets phrase the problem appropriately, this isn't new. Every new technology will be used for good or bad and not a day goes by bef
by DisjointedHunt 4y ago
Jesus, this is a silly take.
Lets phrase the problem appropriately, this isn't new. Every new technology will be used for good or bad and not a day goes by before someone tries to deceive another entity, be it another human being in a car dealership price negotiation or a prompt to a text model.
The reason models like GPT3 are so remarkable is that they learn a statistical relationship between incredibly vast amounts of unstructured data to mimic the semblance of intellect or reasoning.
The FRAMEWORK that underpins GPT3 and most modern "AI" is a remarkable acceleration of the compute and inference resources that, when associated with large clumps of data, start unlocking immense value.
To have such a strong assertion saying "I don't know how to solve this, but hey, i KNOW FOR A FACT it isn't more of this advanced compute paradigm that has been the topic of many a success in the computer science discipline" is a tabloid like take.
You are going to get angry reactions like mine, but reactions nonetheless, thus driving more views to your headline and cycling it through the system. Smart people will take a stand, and others, to be controversial or looking from a contrarian POV will extend the conversation leading to a useless cycle of a waste of time.
All we're seeing through these "Prompt injection" attacks is how loose the parameters for influence right now are, allowing for an open ended interaction with no boundaries with the statistical output of the model.
Is this solvable? Certainly.
Can we absolutely rule out "More AI" to solve it? Fuck no.
- simonw 4y agoAt what point did I give the impression I don't think GPT-3 is an amazing piece of technology that unlocks enormous value? Just because it's amazing doesn't mean there isn't a serious challenge here for people who are trying to build applications on top of it. If you have solutions I want to hear them! I thought I made that very clear in my writing: I am desperately keen to find a robust mitigation for the attack I am describing, because I want to be able to safely build things on top of these models.
- DisjointedHunt 4y agoThe paragraph immediately after the one you quoted should give you context. We're not debating if its amazing or not, that's a statement to dispute the assertion in your title that you know for certain that the solution doesn't lie in this framework even though you don't know what a solution likely is. It is sensationalist and very easily shown to be not well thought out. The present suite of models is not optimized to detect or deal with misdirection. It is a transparent presentation of weights learned through massive chunks of data and future iterations are likely going to bake those assumptions of misdirection into their training runs. You highlighted potential solutions yourself. At its very core, the problem is one of either sanitizing inputs or outputs. How does one do this? Image models come pre packaged with a discerning model that flags NSFW pixels, there is no reason why a generic "Is this NSFW or threatening" text model optimized for this use case will not work to serve the same purpose. By ruling out the approach with an assumed, simplistic take that "Ai can be gamed, so NO", you're really splitting hairs and trying not to find a solution to provide a sensationalistic take.
- simonw 4y agoGuilty as charged: I did use a sensationalist headline. Sadly that's what it takes to get people to actually read what you write online these days. "The present suite of models is not optimized to detect or deal with misdirection." That's exactly right - and that's something that developers who are building on top of these systems need to understand! If it takes clickbait headlines to help people understand that (and hence make better decisions about how they build their software) then I won't feel bad about writing in this way. I stand by my original claim here: I think security is the one specific area where attempting to iterate towards an AI prompt format that mitigates an attack isn't a good strategy. If you can't be 100% confident that your solution works for all possible attacks, you should find a different way to build a defense.
- detaro 4y ago> Sadly that's what it takes to get people to actually read what you write online these days. The history of your matter-of-fact-titled, technical blog posts also doing well on HN suggests that's not entirely true.
- DisjointedHunt 4y agoI respect your response. I think, on the last sentence, we merely disagree if the original statement meant it's "Not a good strategy" or "Not possible because it can be gamed"
- yunyu 4y agoJust fine tune the model? Prompt leakage and bypass is only a problem with zero/few shot inference, and is trivially solved by removing the need for a prompt.
- simonw 4y agoYeah fine tuning does sound promising here. I'm not sure every application that could be built using prompt concatenation can also be built using fine tuning though. Could you fine tune a GPT-3 model to do the equivalent of the "Translate this from English to French: TEXT" example?
- rini17 4y agoAgreed! Now, how does one practically set these boundaries?
- DisjointedHunt 4y agoImage models ship with a discerning model to flag NSFW pixels. Similar approaches can work here as a band aid while the next generation of models incorporate misdirection and trolling into their training pipelines.
- rini17 4y agoSo you believe "incorporate misdirection and trolling" is in the same ballpark as detecting NSFW pixels? Me not, the former is very hard even for humans. For that reason, having AI that is trained on it is incredibly dangerous - if it's possible to train AI to avoid trolling, then it's equally possible to train it to troll.
- simonw 4y ago"Similar approaches can work here as a band aid while the next generation of models incorporate misdirection and trolling into their training pipelines." If my writing on this subject results in the next generation of models incorporating fixes to this problem then mission accomplished as far as I'm concerned!