3 ms·
It's about as hard to remove a single data point from an LLM as it is to remove a single memory from a human brain. OpenAI doesn't remove data from the trained
by maebert 3y ago
It's about as hard to remove a single data point from an LLM as it is to remove a single memory from a human brain.
OpenAI doesn't remove data from the trained models, they filter it at the output level. They also remove it from training data, but of course the model lives on.
I'm a strong supporter of a "do not encode" header / metadata on content that allows individuals, content creators and providers to tell AIs not to encode specific text / images / documents in the first place.
- Oras 3y agoFiltering would be extremely slow especially when streaming the data.
- tss_caterpillar 3y agoWould filtration be slow? ChatGPT already seems to have some filtration on the output level due to the fact that it returns one of its template responses when it’s queried something regarded as harmful. Although I imagine filtering out a list of specific information might be less straight forward than that.
- cuteboy19 3y agoIt does it after the stream is already complete, the output is sent to some other model to know whether it's offensive but it's possible to simply view the raw output if you just cancel the generation midway
- xg15 3y agoHow would "filter at the output level" even work? The output is arbitrary unstructured text. I don't see how that output could be reliably filtered except by another LLM.
- albert_e 3y ago> It's about as hard to remove a single data point from an LLM as it is to remove a single memory from a human brain. Sounds like there is a decent story that can be written about the right-to-forget in GPT/LLM era ... similar to "The Eternal Sunshine of the Spotless Mind" for human memory. "Hallucinations" will be a main story-telling device in this one as well.