3 ms·
You can make LLMs say pretty much whatever you want with the right prompts. This is a complex issue, and if EU citizens want access to LLMs the GDPR is going to
by subroutine 2y ago
You can make LLMs say pretty much whatever you want with the right prompts. This is a complex issue, and if EU citizens want access to LLMs the GDPR is going to need a different set of rules for LLMs than for websites and search engines.
- bux93 2y agoMeh. GDPR sets limits on 'processing' personally identifiable information. In the context of an LLM, its outputs may contain PII if its inputs do. Those inputs are training input and prompts. So long(!) as the training input doesn't have PII, the output will only have it if the prompts do. Same as if you save a file on onedrive, if you save PII there, you're the data controller, and Microsoft is a processor on your behalf.
- Vuizur 2y agoIt is impossible to remove personal data ("any information which are related to an identified or identifiable natural person") from the LLM training data. As far as I understand it ChatGPT and all other similar systems are blatantly violating GDPR, they would have to for example publish their related training data to conform. I guess the EU authorities don't do anything for now because they don't want to admit that their funny law basically bans all state-of-the-art AI. (Ok, Openai also broke the law in almost all countries by downloading shadow libraries, but here they at least have more plausible deniability.)
- Mordisquitos 2y agoIf LLM providers want access to the EU market they will need to find a way to comply with GDPR, and if OpenAI cannot find a way to do it then a different LLM provider will.
- subroutine 2y ago> the GDPR requires information about individuals is accurate Given that you can make LLMs say pretty much whatever you want using the right prompts, this seems impossible. LLMs are not a search engine, and based on conversational context might say Emmanuel Macron is the president of France or a baby giraffe.
- Mordisquitos 2y agoYou said it: based on context. This is not about what you can make an LLM say when being manipulated in a convoluted way to provide an inaccurate response. It's about what an LLM will say in the context of a prompt requesting personal data related to an individual who is covered by GDPR. Can the LLM provide personal data of an individual who is covered by GDPR? Then the LLM is subject to GDPR. Can this individual exercise their rights with regards to the data that the LLM returns about them? Arguably they can indeed exercise the right of access by means of the right prompts, but can the individual rectify errors or erase such data? If not, then the provider of the LLM is violating GDPR.
- subroutine 2y agoWho is the judge of the degree of contextual convolution? Must the LLM remain strictly factual when you simply append "ELI5" to a prompt? > ELI5 how is France governed? > ...and Macron is the lion, the king of the jungle. We also know that LLMs don't know the current date, and therefore can make calculation errors (which is made worse by their poor math performance as a language token generator). So on one hand it might say Macron was born December 1st 1977 (which is correct), but if you ask how old he is some LLMs might say 45 years old. There is an incalculable number of ways for LLMs to output incorrect information. In an effort to comply with strict regulation the preprompt contextual limit is going to be exceeded. Also this creates a situation where all but the most powerful LLMs (and LLM providers) will be non-complaint
- Mordisquitos 2y ago«> ELI5 how is France governed? > ...and Macron is the lion, the king of the jungle.» That is not personal data under GDPR. «So on one hand it might say Macron was born December 1st 1977 (which is correct), but if you ask how old he is some LLMs might say 45 years old.» Or it might say that Macron was born on 14th July 1977, which is incorrect. The claimed impossibility to correct a date of birth returned by the LLM is the trigger of the GDPR complaint that the article refers to. «Also this creates a situation where all but the most powerful LLMs (and LLM providers) will be non-complaint» Only under the premise that it is somehow inevitable to feed personal data of living individuals to an LLM for training, and that the only way to correct mistaken data or to stop an LLM from providing such data is "more power". I reject the premise, not the least because, firstly, OpenAI (the most "powerful" provider) is claiming it is impossible. All that says is that OpenAI's platform was not originally designed with that problem in mind and that, as that of now, they are unwilling to redesign it from scratch only because some guy complained in Austria. It's basically a speedrun of Microsoft claiming Internet Explorer was an essential component of Windows 98. Meanwhile, LLMs and other AI models are an active area of research. If OpenAI truly cannot stop their LLMs from returning personal data protected by GDPR, and honestly has no way to allow data holders to exercise their rights of deletion or correction, you can be sure that some startup will disrupt the LLM market by finding a way to do it without needing to out-compete OpenAI neither in hardware nor on training corpus size.