3 ms·
Metadata is very low hanging fruit for document watermarking. Typically the PDF renderer will use spacing, kerning, invisible characters, and all sorts of stega
by nonrandomstring 2y ago
Metadata is very low hanging fruit for document watermarking.
Typically the PDF renderer will use spacing, kerning, invisible
characters, and all sorts of steganography to make each copy unique.
What would be the point of a hash? More likely the hash is a MAC,
that's been salted with some secret plus the unique copy. That would
help the publisher identify a laundered copy. With two or more copies
its possible to re-anonymise. That's actually something I wonder
whether summarising language models would be good at. Of course they
may also make steganographic alterations to diagrams.
Because PDFs are such dirty documents I almost always convert them to
plain text, usually with no loss of semantics.
- amelius 2y agoYes, but don't give them any ideas.
- nicce 2y agoEven plain text has its risks: https://news.ycombinator.com/item?id=33621562 https://news.ycombinator.com/item?id=33621562
- sodality2 2y agoIf they modify academic papers’ actual source content they would lose all of their trust. Imagine they doctor some data just to create more diffs
- nicce 2y agoThey modify it all the time by providing different access formats and with editors. But it does not change the meaning. Whitespaces do not alter the meaning.