7 ms·
New information extracted from Snowden PDFs through metadata version analysis
- pfisherman 9mo agoCan someone spell out how this is possible? Do pdfs store a complete document version history? Do they store diffs in the metadata? Does this happen each time the document is edited?
- alhirzel 9mo agoPDFs are just a table of objects and tree of references to those objects; probably, prior versions of the document were expressed in objects with no references or something like that.
- flotzam 9mo agoAt the bottom of the page there's a link to the pdfresurrect package, whose description says "The PDF format allows for previous changes to be retained in a revised version of the document, thereby keeping a running history of revisions to the document. This tool extracts all previous revisions while also producing a summary of changes between revisions."
- QuantumNomad_ 9mo agoNeat! https://github.com/enferex/pdfresurrect https://github.com/enferex/pdfresurrect
- aidos 9mo agoYou can replace objects in PDF documents. A PDF is mostly just a bunch of objects of different types so the readers know what to do with them. Each object has a numbered ID. I recommend mutool for decompressing the PDF so you can read it in a text editor: mutool clean -d in.pdf out.pdf If you look below you can see a Pages list (1 0 obj) that references (2 0 R) a Page (2 0 obj). 1 0 obj << /Type /Pages /Count 1 /Kids [ 2 0 R ] >> endobj 2 0 obj << /Type /Page /Contents 5 0 R ... >> endobj Rather than editing the PDFs in place, it's possible to update these objects to overwrite them by appending a new "generation" of an object. Notice the 0 has been incremented to a 1 here. This allows leaving the original PDF intact while making edits. 1 1 obj << /Type /Pages /Count 2 /Kids [ 2 0 R 200 0 R ] >> endobj You can have anything inside a PDF that you want really and it could be orphaned so a PDF reader never picks up on it. There's nothing to say an object needs to be referenced (oh, there's a "trailer" at the end of the PDF that says where the Root node is, so they know where to start).
- SeriousM 9mo agoTo put it reaaaaaly simple, a PDF is like a notion document (blocks and bricks) with a git-like object graph?
- aidos 9mo agoHa! As if anything about Notion is simple. But yeah. It's all just objects pointing at each other. It's mostly tree structured, but not entirely. You have a Catalog of Pages that have Resources, like Fonts (that are likely to be shared by multiple pages hence, not a tree). Each Page has Contents that are a stream of drawing instructions. This gives you a sense of what it all looks like. The contents of a page is a stack based vector drawing system. Squint a little (or stick it through an LLM) and you'll see Tf switches to Font F4 from the resources at size 14.66, Tj is placing a char at a position etc. 2 0 obj << /Type /Page /Resources << /Font << /F4 4 0 R >> >> /Contents 5 0 R >> endobj 5 0 obj << /Length 340 >> stream q BT /F4 14.66 Tf 1 0 0 -1 0 .47981739 Tm 0 -13.2773438 Td <002B> Tj 10.5842743 0 Td <004C> Tj ET Q... endstream endobj I'm going to hand wave away the 100+ different types of objects. But at it's core it's a simple model.
- pfisherman 9mo agoThanks for the technical explanation! This is pretty fascinating. So it works kind of like a soft delete — dereference instead of scrubbing the bits. Is this behavior generally explicitly defined in PDF editors (i.e. an intended feature)? Is it defined in some standard or set of best practices? Or is it a hack (or half baked feature) someone implemented years ago that has just kind of stuck around and propagated?
- clord 9mo agoThe intention is to make editing easy and quick on slow and memory deficient computers. This is how for example editing a pdf with form field values can be so fast. It’s just appending new values for those nodes. If you need to omit edits you’d have to regenerate a fresh pdf from the root.
- 1317 9mo agohttps://hackerfactor.com/blog/index.php?/archives/1085-A-Typical-PDF.html https://hackerfactor.com/blog/index.php?/archives/1085-A-Typ...
- alhirzel 9mo agoThere needs to be better tooling for inspecting PDF documents. Right now, my needs are met by using `qpdf` to export QDF [1], but it is just begging for a GUI to wrap around it... [1] https://qpdf.readthedocs.io/en/stable/qdf.html https://qpdf.readthedocs.io/en/stable/qdf.html
- soared 9mo agoIn what contest do you use that tool? Looks like that page is primarily about editing pdfs using that format rather than inspecting. Very tempting to fool around with the ideas especially after the Epstein pdf debacle.
- deleted 9mo ago[deleted]
- piffey 9mo agoTake a look at the REMNux reverse engineering page for PDF documents (https://docs.remnux.org/discover-the-tools/analyze+documents/pdf https://docs.remnux.org/discover-the-tools/analyze+documents...). Lots of tools here for looking at malicious PDFs that can be used to inspect/understand even non-malicious documents.
- treetalker 9mo ago% pdfresurrect -w epsteinfiles.pdf
- cypherpunks01 9mo agoAnyone tried this?
- huflungdung 9mo ago[dead]
- snek_case 9mo agoWeekend project?
- jokoon 9mo agoI have read claims that there were fake documents inserted in those leaks, who aimed at pushing disinformation.
- Ms-J 9mo agoThis is insightful work, great job. Recently someone else revisited the Snowden documents and also found more info, but I can't recall the exact details. Snowden and the archives were absolute gifts to us all. It's a shame he didn't release everything in full though.
- dfreel 9mo ago[flagged]
- zoklet-enjoyer 9mo agoSpy on me harder, daddy
- Andrex 9mo agoThe best way to fix a problem is to bring it into the light, not pretend it doesn't exist. "Security by obscurity" has been debunked for decades. If our system is so flawed Snowden's leaks would have blown everything up, maybe the system deserves to be blown up. Otherwise we're just papering over flaws which likely will be discovered and exploited eventually.
- dfreel 9mo ago[flagged]
- wizzwizz4 9mo agoAre you thinking of Julian Assange? I'm not aware of Edward Snowden releasing anything that might have resulted in the deaths of field agents.
- threethirtytwo 9mo agoThat’s a one sided view. Secrecy is also used to hide corruption and crimes. There is plenty of corruption going on just within the CIA.
- 9mo ago
- password4321 9mo agoThe "print and scan physical papers back to a PDF of images" technique for final release is looking better and better from an information protection perspective.
- cookiengineer 9mo ago> The "print and scan physical papers back to a PDF of images" technique for final release is looking better and better from an information protection perspective. Note that all (edit: color-/ink-) printers have "invisible to the human eye" yellow dotcodes, which contain their serial number, and in some cases even the public IP address when they've already connected to the internet (looking at you, HP and Canon). So I'd be careful to use a printer of any kind if you're not in control of the printer's firmware. There's lots of tools that started to decode the information hidden in dotcodes, in case you're interested [1] [2] [3] [1] https://github.com/Natounet/YellowDotDecode https://github.com/Natounet/YellowDotDecode [2] https://github.com/mcandre/dotsecrets https://github.com/mcandre/dotsecrets [3] (when I first found out about it in 2007) https://fahrplan.events.ccc.de/camp/2007/Fahrplan/events/1976.en.html https://fahrplan.events.ccc.de/camp/2007/Fahrplan/events/197...
- everdrive 9mo ago>Note that all printers have "invisible to the human eye" yellow dotcodes, which contain their serial number, and in some cases even the public IP address when they've already connected to the internet (looking at you, HP and Canon). I've got a black and white brother printer which uses toner. Is there something similar for this printer?
- gramie 9mo agoI believe that this only exists for colour printers. The official reasoning was to trace people counterfeiting money.
- cookiengineer 9mo ago> black and white brother printer excellent choice, that's what I am using. Also it's Linux / CUPS compatible and without a broken proprietary rasterizer.
- londons_explore 9mo agoSo this is almost certainly redaction by the journalists? It is disappointing they didn't mark those sections "redacted", with an explanation of why. It is also disappointing they didn't have enough technical knowhow to at least take a screenshot and publish that rather than the original PDF which presumably still contains all kinds of info in the metadata.
- libroot 9mo agoYes, the journalists did the redactions. The metadata timestamps in one of the documents show that the versions were created three weeks before the publication. And to be honest, the journalists generally have done a great work on pretty much in all the other published PDFs. We've went through hundreds and hundreds of the published documents, and these two documents were pretty much the only ones which had metadata leak by a mistake revealing something significant (there are other documents as well with metadata leaks/failed redactions, but nothing huge). Our next part will be a technical deep-dive on PDF forensic/metadata analysis we've done.
- DANmode 9mo agoGreat work, great comment. Thank you.
- layer8 9mo agoThese PDFs apparently used the “incremental update” feature of PDF, where edits to the document are merely appended to the original file. It’s easy to extract the earlier versions, for example with a plain text editor. Just search for lines starting with “%%EOF”, and truncate the file after that line. Voila, the resulting file is the respective earlier PDF version. (One exception is the first %%EOF in a so-called linearized PDF, which marks a pseudo-revision that is only there for technical reasons and isn’t a valid PDF file by itself.)
- ajross 9mo agoIt's hilarious the extent to which Adobe Systems's ridiculously futile attempt to chase MS Word features ended up being the single most productive espionage tool of the last quarter century.
- layer8 9mo agoI don’t think this was particularly modeled on MS Word. The incremental update feature was introduced with PDF 1.2 in 1996. It allows to quickly save changes without having to rewrite the whole file, for example when annotating a PDF. Incremental updates are also essential for PDF signatures, since when you add a subsequent signature to a PDF, you couldn’t rewrite the file without breaking previous signatures. Hence signatures are appended as incremental updates.
- cubefox 9mo agoI'm pretty sure you can change various file formats without rewriting the entire file and without using "incremental updates".
- deleted 9mo ago[deleted]
- mattzito 9mo agoNo, if you are going to change the structure of a structured document that has been saved to disk, your options are: 1) Rewrite the file to disk 2) Append the new data/metadata to the end of the existing file I suppose you could pre-pad documents with empty blocks and then go modify those in situ by binary editing the file, but that sounds like a nightmare.
- c-c-c-c-c 9mo ago> We contacted Ryan Gallagher, the journalist who led both investigations, to ask about the editorial decision to remove these sections. After more than a week, we have not received a response. Hopefully we'll hear something now that the Christmas holidays are over.
- echelon 9mo agoWhy are the journalists redacting the docs? That's incredibly puzzling. Is there something in here so damaging that they refuse to publish it? Did the government tell them they'd be in trouble if they published it? Are the journalists the only ones with access to the raw files?
- nacozarina 9mo agoTraditionally an editor would be obligated to review the material and redact info that could be harmful to others. The publisher has distinct liability independent of govt opinion.
- mlmonkey 9mo ago> and redact info that could be harmful to others. of course, these concerns are only applicable when these "others" are Americans and the American institutions. Everybody else can just fend for themselves. Whats good for the goose, should be good for the gander. If American journalists feel like there is no problem with disclosing secrets of, say, Maduro, then they should not be protecting people like Trump (just as an example).
- tolerance 9mo ago[flagged]
- rendx 9mo agoAre you asking how much was done with pen and paper, and how much of it was done on a computer, i.e. machine assisted? Where do you draw the line? How is "hands-on" in contrast to anything? Is it only "hands-on" when you don't use any tool to assist you? I suspect you're inquiring about the use of LLMs, and about that I wonder: Why does it matter? Why are you asking?
- Ylpertnodi 9mo ago[flagged]
- rendx 9mo agoAre you confusing me with the authors, or why would you think I could? And I'm asking 'tolerance' to clarify their question, which means I wouldn't be able to answer it even if I had the knowledge they were after, since I don't understand what they're asking.
- tolerance 9mo agoFirst thanks for taking my question seriously and not as just a rib and asking a lot of questions in return that I want to consider myself. By "hands-on" I'm asking whether the provided insight is the product of human intellection. Experienced, capable and qualified. Or at least an earnest attempt at thinking about something and explaining the discoveries in the ways that thinking was done before ChatGPT. For some reason I find myself using phrases involving the hands (etc. hands-on, handmade, hand-spun) as a metaphor for work done without the use of LLMs. I emphasize insight because I feel like the series of work on the Snowden documents by libroot is wanting in that. I expressed as much the last time their writing hit the front page: <https://news.ycombinator.com/item?id=46236672 https://news.ycombinator.com/item?id=46236672>. These are summaries. I don't think that it yields information that can't otherwise be pointed out and made mention of by others; presumably known and reputable. With as high-profile of an event that this is I'd expect someone covering it almost 16 years later to tell us beyond what when judged on the merit of its import amounts to a motivated section of the ‘Snowden disclosures’ Wikipedia entry. The discussion that this series invites typically is centered around people's thoughts about the story of the Snowden documents in general, and in this case exchanges about technical aspects like how PDF documents work and can be manipulated in general. The one comment that I feel addresses the actual tension embedded in the article—"Who edited the documents?"—leads to accusations that the documents were tampered with by the media: <https://news.ycombinator.com/item?id=46566372 https://news.ycombinator.com/item?id=46566372>. I don't think that that's an implausible claim but I find issue with it being made with such confidence by the anonymous source behind the investigations (I'm withholding ironically putting "investigations" in...nevermind). If the author actually provided something that advanced to the reader why this information is significant, what to do with or think about it and how they came about discovering the answers to the aforementioned 'why' and ‘what’ and additionally why they’re word ought to matter to us at all, I'd be less inclined to speculate that this is just someone vibe sleuthing their way through documents that on the surface are only significant to the public as the claim "the government is spying on you" is. This particular post uncovers some nice information. It's a great find. I'm in no position to investigate whether it was already known. But what are we supposed to learn from it aside from "one of the documents were changed before it was made public". What's significant about the redaction? Is Ryan Gallagher responsible? Or does he know who is. Is he at all obliged to explain this to a presumably anonymous inquirer? Or is it now the duty of the public to expect an explanation as affected by said anonymous inquirer? Remember when believing that the government was rife with pedophiles automatically associated you with horn-helmet-wearing insurrectionists?
- pseudosavant 9mo agoIn addition to the print paper and scan approach, I do wonder how effective it would be to “Print to XPS” and then “print” that into a PDF.
- bawolff 9mo agoIts crazy this is just being discovered now.
- farceSpherule 9mo ago[dead]