3 ms·
Just so people are clear, these types of models are almost universally naive and basic. If all you have is a single generic neutral message, "Hi, this is Bob."
by CMay 6mo ago
Just so people are clear, these types of models are almost universally naive and basic. If all you have is a single generic neutral message, "Hi, this is Bob.", it will be sufficient in most cases. If you have a pile of data, I am not aware of any PII redaction tool that has factored in all of the risks to identity leakage.
The problem is when companies use things like this and somehow believe they are anonymizing the data. No, you are not.
Still, for scenarios where the processed data isn't being directly published or shared, but used as some intermediate step like moderation enforcement, human evaluation layers or model training it can be useful to filter these things out.
- deleted 6mo ago[deleted]