2 ms·
>Comments should be passed to the model with clear role boundaries that prevent them from being interpreted as system-level directives. Well, such clear bounda
by wrs 3mo ago
>Comments should be passed to the model with clear role boundaries that prevent them from being interpreted as system-level directives.
Well, such clear boundaries would solve lots of problems. But those don’t exist, do they?
- InsideOutSanta 3mo agoYeah, I suspect the main reason this was rejected is simply because it's not fixable. This is just how LLMs work. This LLM ingests untrusted data, so there will always be a non-zero chance that this type of prompt injection succeeds.
- 27183 3mo agoThis makes me crazy. When I started my career in software the focus was on security, correctness, uptime (measured in nines--remember those?), and performance. Features are important, but building a crap feature is worse than building no feature at all. I don't understand how these systems are passing the bar. You would have been fired for trying to railroad something like this into production 10-15 years ago. What happened?
- wrs 3mo agoI suspect when those 10-digit wire transfers start arriving in your bank account, your attitude changes rapidly.
- 27183 3mo agoSpeaking personally, approximately one minute after a 10 digit wire transfer arrives in my account I will disappear permanently to sail the seas on my yacht. What would be the incentive to continue working? Personal finances aside, that's no way to run a business. Torching your brand, alienating your users, and pissing off your customers is a well known path to ruin. [edit] Even if it results in some temporary windfall--is the thesis that the windfall will be so big they no longer need users or customers? It eludes me completely what the companies that are building this trash now are hoping to achieve. It's especially galling that publicly traded companies are doing it. It's one thing for a startup to blow a bunch of venture capital on a speculative, half-baked product idea. Great risks sometimes yield great rewards. Usually they don't. Founders and VCs knowingly and willingly sign up for those risks. It's a totally different story for public companies.
- chias 3mo agoAh yes - the cure for world hunger: eating food.
- mattalex 3mo agoYou can get rid of 99.9% of those attacks by simply dispatching the data consumption to a different instance of the LLM, see, for instance, some of the later patterns in https://arxiv.org/abs/2506.08837 https://arxiv.org/abs/2506.08837
- iqihs 3mo agoThanks for the article link! Do you happen to know where to follow/read more articles like this for someone interested in getting more into AI security? Ty
- g-b-r 3mo agoHow would they apply to this case? They require being able to transorm the output to something symbolic, but this YouTube feature necessarily has to output free-form text, derived directly from the comments..! What would actually prevent the "attack" is for YouTube to not turn markdown from random LLM outputs into actual links. In general, those patterns seem applicable only to a limited amount of cases, I think that they prevent much less than 99.9% of the attacks.