4 ms·
(I lead privacy at Brave and am one of the authors) > Instead they believe model alignment, trying to understand when a user is doing a dangerous task, etc. wi
by skaul 1y ago
(I lead privacy at Brave and am one of the authors)
> Instead they believe model alignment, trying to understand when a user is doing a dangerous task, etc. will be enough.
No, we never claimed or believe that those will be enough. Those are just easy things that browser vendors should be doing, and would have prevented this simple attack. These are necessary, not sufficient.
- cowboylowrez 1y agowhat you're saying is that the described step, "model alignment" is necessary even though it will fail a percentage of the time. whenever I see something that is "necessary" but doesn't have like a dozen 9's for reliability against failure or something well lets make that not necessary then. whadya say?
- skaul 1y agoThat's not how defense-in-depth works. If a security mitigation catches 90% of the "easy" attacks, that's worth doing, especially when trying to give users an extremely powerful capability. It just shouldn't be the only security measure you're taking.
- cowboylowrez 1y agosure sure, except llms. I mean its valid and all bringing up tried and true maxims that we all should know regarding software, but whens the last time the ssl guys were happy with a fix that "has a chance of working, but a chance of not working." defense in depth is to prevent one layer failure from getting to the next, you know, exploit chains etc. Failure in a layer is a failure, not statistically expected behavior. we fix bugs. what we need to do is treat llms as COMPLETELY UNTRUSTED user input as has been pointed out here and elsewhere time and again. you reply to me like I need to be lectured, so consider me a dumb student in your security class. what am I missing here?
- jrflowers 1y ago> what am I missing here? Yeah the tone of that response seems unnecessarily smug. “I’m working on removing your front door and I’m designing a really good ‘no trespassing’ sign. Only a simpleton would question my reasoning on this issue”
- ModernMech 1y ago> what am I missing here? I guess what I don't understand is that failure is always expected because nothing is perfect, so why isn't the chance of failure modeled and accounted for? Obviously you fix bugs, but how many more bugs are in there you haven't fixed? To me, "we fix bugs" sounds the same as "we ship systems with unknown vulnerabilities". What's the difference between a purportedly "secure" feature with unknown, unpatched bugs; and an admittedly insecure feature whose failure modes are accounted for through system design taking that insecurity into account, rather than pretending all is well until there's a problem that surfaces due to unknown exploits?
- cowboylowrez 1y agoI think you're correct with accounting for the security "attributes" of these llms if you're going to use them, like you said, "taking that insecurity into account". If we sit down and examine the statistics of bugs, the costs of their occurance in production and weighed everything with some reasonable criteria, I think we could somehow arrive at a reasonable level of confidence that allows us to ship a system to production. Some organizations do better with this than others of course. During a projects development cycle, we could watch out for common patterns, buffer overflows, use after free for c folks, sql injection or non escaping stuff in web programming but we know these are mistakes and we want to fix them. With llms the mitigation that I'm seeing is that we reduce the errors 90 percent, but this is not a mitigation unless we also detect and prevent the other 10 percent. Its just much more straightforward to treat llms as untrusted, because they are, you're getting input from randos by virtue of its training data. producing mistaken output is not actually a bug, its actually expected behavior, unless you also believe in the tooth fairy lol >To me, "we fix bugs" sounds the same as "we ship systems with unknown vulnerabilities". to me, they sound different ;)
- MattPalmer1086 1y agoDefence in depth means you have more than one security control. But the LLM cannot be regarded as a security control in the first place; it's the thing you are trying to defend against. If you tried to cast an unreliable insider as part of your defence in depth strategy (because they aren't totally unreliable), you would be laughed out of the room in any security group I've ever worked with.
- cowboylowrez 1y agocall it "vibe security" lol
- MattPalmer1086 1y agoHaha, like it!
- kbrkbr 1y agoI am sure that's what you mean, but I think it is important to state it explicitly every now and then: > Defence in depth means you have more than one security control that overlap. Having them strictly parallel is not defense in depth (e.g. on one door to the same room a dog, and on a different unconnected door a guard).
- MattPalmer1086 1y agoYes, fully agree. Should have made that explicit. And also different types of control too. So you might have a lock on the door, a dog, and a pressure sensor on the floor after it...
- ec109685 1y agoThis statement on your post seems to say it would definitively prevent this class of attacks: “In our analysis, we came up with the following strategies which could have prevented attacks of this nature. We’ll discuss this topic more fully in the next blog post in this series.”
- jrflowers 1y agoBut you don’t think that, fundamentally, giving software that can hallucinate the ability to use your credit card to buy plane tickets, is a bad idea? It kind of seems like the only way to make sure a model doesn’t get exploited and empty somebody’s bank account would be “We’re not building that feature at all. Agentic AI stuff is fundamentally incompatible with sensible security policies and practices, so we are not putting it in our software in any way”
- petralithic 1y agoTheir point was that no amount of statistical mitigation is enough, the only way to win the game is to not play, ie not build the thing you're trying to build. But of course, I imagine Brave has invested to some significant extent in this, therefore you have to make this work by whatever means, according to your executives.