3 ms·
This is not even a new problem... Back in 2011, Google faced the same problem mining bi-texts from the Internet for their statistical machine translation softw
by jabbany 4y ago
This is not even a new problem...
Back in 2011, Google faced the same problem mining bi-texts from the Internet for their statistical machine translation software. The thought was that one could utilize things like multi-lingual websites to learn corresponding translations.
They quickly realized that a lot of sites were actually using Google Translate without human intervention to make multi-lingual versions of their site, so naive approaches would cause the model to get trained on its own suboptimal output.
So they came up with a whole watermarking system so that the model could recognize its own output with some statistical level of certainty, and avoid it. It wouldn't be surprising if this is being done for LLMs too. The more concerning problem is when different LLMs, who are not aware of each others' watermarks, end up potentially becoming inbred should the ratio of LLM content rise dramatically...
Ref: https://aclanthology.org/D11-1126.pdf https://aclanthology.org/D11-1126.pdf