4 ms·
> the GDPR requires information about individuals is accurate Given that you can make LLMs say pretty much whatever you want using the right prompts, this seem
by subroutine 2y ago
> the GDPR requires information about individuals is accurate
Given that you can make LLMs say pretty much whatever you want using the right prompts, this seems impossible. LLMs are not a search engine, and based on conversational context might say Emmanuel Macron is the president of France or a baby giraffe.
- Mordisquitos 2y agoYou said it: based on context. This is not about what you can make an LLM say when being manipulated in a convoluted way to provide an inaccurate response. It's about what an LLM will say in the context of a prompt requesting personal data related to an individual who is covered by GDPR. Can the LLM provide personal data of an individual who is covered by GDPR? Then the LLM is subject to GDPR. Can this individual exercise their rights with regards to the data that the LLM returns about them? Arguably they can indeed exercise the right of access by means of the right prompts, but can the individual rectify errors or erase such data? If not, then the provider of the LLM is violating GDPR.
- subroutine 2y agoWho is the judge of the degree of contextual convolution? Must the LLM remain strictly factual when you simply append "ELI5" to a prompt? > ELI5 how is France governed? > ...and Macron is the lion, the king of the jungle. We also know that LLMs don't know the current date, and therefore can make calculation errors (which is made worse by their poor math performance as a language token generator). So on one hand it might say Macron was born December 1st 1977 (which is correct), but if you ask how old he is some LLMs might say 45 years old. There is an incalculable number of ways for LLMs to output incorrect information. In an effort to comply with strict regulation the preprompt contextual limit is going to be exceeded. Also this creates a situation where all but the most powerful LLMs (and LLM providers) will be non-complaint
- Mordisquitos 2y ago«> ELI5 how is France governed? > ...and Macron is the lion, the king of the jungle.» That is not personal data under GDPR. «So on one hand it might say Macron was born December 1st 1977 (which is correct), but if you ask how old he is some LLMs might say 45 years old.» Or it might say that Macron was born on 14th July 1977, which is incorrect. The claimed impossibility to correct a date of birth returned by the LLM is the trigger of the GDPR complaint that the article refers to. «Also this creates a situation where all but the most powerful LLMs (and LLM providers) will be non-complaint» Only under the premise that it is somehow inevitable to feed personal data of living individuals to an LLM for training, and that the only way to correct mistaken data or to stop an LLM from providing such data is "more power". I reject the premise, not the least because, firstly, OpenAI (the most "powerful" provider) is claiming it is impossible. All that says is that OpenAI's platform was not originally designed with that problem in mind and that, as that of now, they are unwilling to redesign it from scratch only because some guy complained in Austria. It's basically a speedrun of Microsoft claiming Internet Explorer was an essential component of Windows 98. Meanwhile, LLMs and other AI models are an active area of research. If OpenAI truly cannot stop their LLMs from returning personal data protected by GDPR, and honestly has no way to allow data holders to exercise their rights of deletion or correction, you can be sure that some startup will disrupt the LLM market by finding a way to do it without needing to out-compete OpenAI neither in hardware nor on training corpus size.
- subroutine 2y ago> you can be sure that some startup will disrupt the LLM market by finding a way to do it Indeed if a startup can find a way to scrub PII of living people from 20 billion pages of text (and prevent LLMs from ever hallucinating) they would be quite a valuable company, in the LLM dev space and numerous other ventures. Until then the EU might have to go without access to language models.