3 ms·
The paragraph immediately after the one you quoted should give you context. We're not debating if its amazing or not, that's a statement to dispute the assertio
by DisjointedHunt 4y ago
The paragraph immediately after the one you quoted should give you context. We're not debating if its amazing or not, that's a statement to dispute the assertion in your title that you know for certain that the solution doesn't lie in this framework even though you don't know what a solution likely is.
It is sensationalist and very easily shown to be not well thought out.
The present suite of models is not optimized to detect or deal with misdirection. It is a transparent presentation of weights learned through massive chunks of data and future iterations are likely going to bake those assumptions of misdirection into their training runs.
You highlighted potential solutions yourself. At its very core, the problem is one of either sanitizing inputs or outputs. How does one do this? Image models come pre packaged with a discerning model that flags NSFW pixels, there is no reason why a generic "Is this NSFW or threatening" text model optimized for this use case will not work to serve the same purpose. By ruling out the approach with an assumed, simplistic take that "Ai can be gamed, so NO", you're really splitting hairs and trying not to find a solution to provide a sensationalistic take.
- simonw 4y agoGuilty as charged: I did use a sensationalist headline. Sadly that's what it takes to get people to actually read what you write online these days. "The present suite of models is not optimized to detect or deal with misdirection." That's exactly right - and that's something that developers who are building on top of these systems need to understand! If it takes clickbait headlines to help people understand that (and hence make better decisions about how they build their software) then I won't feel bad about writing in this way. I stand by my original claim here: I think security is the one specific area where attempting to iterate towards an AI prompt format that mitigates an attack isn't a good strategy. If you can't be 100% confident that your solution works for all possible attacks, you should find a different way to build a defense.
- detaro 4y ago> Sadly that's what it takes to get people to actually read what you write online these days. The history of your matter-of-fact-titled, technical blog posts also doing well on HN suggests that's not entirely true.
- DisjointedHunt 4y agoI respect your response. I think, on the last sentence, we merely disagree if the original statement meant it's "Not a good strategy" or "Not possible because it can be gamed"