4 ms·
https://en.wikipedia.org/wiki/K-anonymity https://en.wikipedia.org/wiki/K-anonymity Basically keeping only chunks of text which have are not unique and contribu
by 65a 3y ago
https://en.wikipedia.org/wiki/K-anonymity https://en.wikipedia.org/wiki/K-anonymity Basically keeping only chunks of text which have are not unique and contributed by many users.
- bastawhiz 3y ago> Given person-specific field-structured data That's not what this data is and that's broadly not what LLMs are trained on.
- 65a 3y agoBy chunking and binning the data by user contribution count, you are structuring it as described. Pretraining is often accomplished on similar chunked text. k-anonymity is not perfect though, or state of the art, the wikipedia explains several known attacks.