4 ms·
I wonder if a simple model trained only to spot and report on suspicious injection attempts, or otherwise review the "long-term memory" could be used in the pip
by taberiand 2y ago
I wonder if a simple model trained only to spot and report on suspicious injection attempts, or otherwise review the "long-term memory" could be used in the pipeline?
- hibikir 2y agoSome will have to be built, but the attackers will also work on beating them. It's not like the malicious side of SEO, trying to sneak malware into ad networks, or bypassing a payment processor's attempts at catching fraudulent merchants. A traditional red queen game. What makes this difficult is that the traditional constraints to the problem that provide advantage to the defender in some of those questions (like the payment processor) are unlikely to be there in generative AI, as it might not even be easy to know who is poisoning your data, and how they are doing it. By reading the entire internet, we are inviting in all the malicious content in, as being cautious also makes the model worse in other ways. It's going to be trouble. Out only hope is that economically viable poisoning of the AI's outputs doesn't become economically viable. Incentives matter: See how ransomware flourished when it became easier to get paid. Or how much effort people will dedicate to convincing VCs that their basically fraudulent startup is going to be the wave of the future. So if there's hundreds of millions of dollars in profit from messing with AI results, expect a similar amount to be spent trying to defeat every single countermeasure you will imagine. It's how it always works.
- dijksterhuis 2y ago> So if there's hundreds of millions of dollars in profit from messing with AI results, expect a similar amount to be spent trying to defeat every single countermeasure you will imagine. It's how it always works. Unfortunately that’s not how it has worked in machine learning security. Generally speaking (and this is very general and overly broad), it has always been easier to attack than defend (financially and effort wise). Defenders end up spending a lot more than attackers for robust defences, I.e. not just filtering out phrases. And, right now, there are probably way more attackers. Caveat — been out of the MLSec game for a bit. Not up with SotA. But we’re clearly still not there yet.
- paulv 2y agoIs this not the same as the halting problem (genuinely asking)?
- TZubiri 2y ago[flagged]
- explodes 2y agohttps://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- Tepix 2y agoSounds like Llama guard: https://medium.com/pondhouse-data/llm-safety-with-llama-guard-2-2ce68d6928c5 https://medium.com/pondhouse-data/llm-safety-with-llama-guar...