8 ms·
Up next: 1. Generate a bunch of responses with both Claude and various non-Claude LLMs (ChatGPT, Gemini, Kimi) 2. Train a discriminator model that can differe
by qurren 24d ago
Up next:
1. Generate a bunch of responses with both Claude and various non-Claude LLMs (ChatGPT, Gemini, Kimi)
2. Train a discriminator model that can differentiate Claude vs. non-Claude
3. Train a de-watermarking model using the discriminator model as loss
- hughw 24d agoOr just write a 6 line program to remove the metainfo from the file?
- aero-glide2 24d agovery difficult to remove text watermarking
- ProfessorLayton 24d agoOutputs can be easily translated from the watermarked language to another, either by just knowing two+ languages, or via non-LLM translation.
- HoldOnAMinute 24d agoThere must be some set of outputs that is completely resistant to watermarking. For example if I told Claude to output "Hello, World"
- rcktmrtn 24d agoYour observation is sharp, but it's not just metainfo—it's load-bearing text. To remove it, you need to delve deep and alter the tapestry of carefully selected words. edit: almost forgot the "—"
- Lockal 23d agos/text/prose/ s/'/’/ The honest seam is obvious: claudish works; imitation does not. That’s not nothing.
- DonsDiscountGas 23d agoYou just described how Pangram works.