3 ms·
I also don’t think these companies are lying at all, but I definitely think they’re training on all your data, toggle or not. It’s truly trivial to “anonymize”
by preg_match 19d ago
I also don’t think these companies are lying at all, but I definitely think they’re training on all your data, toggle or not.
It’s truly trivial to “anonymize” and distill your prompts and model output. They could use just about any off-the-shelf cheap model for this. In fact, their TOS explicitly allows this, even with the toggle checked.
What that probably means is that the EXACT content of your prompt is secret. But the actual ideas are not. If you discover something truly novel, then yeah they get that. They can absorb trends in consumer behavior, too.
I’m sure if someone had access to all my paraphrased prompts, which retain 0% of my exact wording, they could find out literally everything about me. It’s a bit like how collecting metadata is as good (or better!) than collecting the real data.
And we all know “anonymizing” data doesn’t really exist like we think it does. Just removing names and identifiers doesn’t make anything anonymous for motivated actors. Or… say… an LLM that is trained to recognize patterns in text. Which is, like, all of them.