9 ms·
Recreating Epstein PDFs from raw encoded attachments
- linuxguy2 8mo agoLove this, absolutely looking forward to some results.
- FarmerPotato 8mo agoIf only Base64 had used a checksum.
- iwontberude 8mo agoThis one is irresistible to play with. Indeed a nerd snipe.
- netsharc 8mo agoI doubt the PDF would be very interesting. There are enough clues in the human-readable parts: it's an invite to a benefit event in New York (filename calls it DBC12) that's scheduled on December 10, 2012, 8pm... Good old-fashioned searching could probably uncover what DBC12 was, although maybe not, it probably wasn't a public event. The recipient is also named in there...
- RajT88 8mo agoThere's potentially a lot of files attached and printed out in this fashion. The search on the DOJ website (which we shouldn't trust), given the query: "Content-Type: application/pdf; name=", yields maybe a half dozen or so similarly printed BASE64 attachments. There's probably lots of images as well attached in the same way (probably mostly junk). I deleted all my archived copies recently once I learned about how not-quite-redacted they were. I will leave that exercise to someone else.
- notenlish 8mo agoThere's 70 results that come out when searching for "application/pdf" on the doj website
- netsharc 8mo agoOK, but if the solution is to brute-force them, there's probably a need to choose which files to focus on. Of course there are other content-types, e.g. searching for "Content-Type: image/jpeg" gets hits as well. But only a few of them actually have the base64 data, mostly there are just the MIME headers.. Looking for "/9j/" (which is Base64 for FF D8 FF, which is the header for JPEG files), the Trumpian justice.gov website ignores "/" and shows results case-insensitively, but there are 4 or 5 base64'ed JPEG images in there. I also saw that the page is vulnerable to code injection, somehow garbage in one search result preview was OCREd as "<s [lots of garbage]>", and the rest of the search results were striken-through because "<s>" is the HTML to do that.
- pimlottc 8mo agoWhy not just try every permutation of (1,l)? Let’s see, 76 pages, approx 69 lines per page, say there’s one instance of [1l] per line, that’s only… uh… 2^5244 possibilities… Hmm. Anyone got some spare CPU time?
- deleted 8mo ago[deleted]
- wahern 8mo agoIt should be much easier than that. You should should be able to serially test if each edit decodes to a sane PDF structure, reducing the cost similar to how you can crack passwords when the server doesn't use a constant-time memcmp. Are PDFs typically compressed by default? If so that makes it even easier given built-in checksums. But it's just not something you can do by throwing data at existing tools. You'll need to build a testing harness with instrumentation deep in the bowels of the decoders. This kind of work is the polar opposite of what AI code generators or naive scripting can accomplish.
- cluckindan 8mo agoOn the contrary, that kind of one-off tooling seems a great fit for AI. Just specify the desired inputs, outputs and behavior as accurately as possible.
- m000 8mo agoYou might be taking the "I" in AI too literally.
- pimlottc 8mo agoI wonder if you could leverage some of the fuzzing frameworks tools like Jepsen rely on. I’m sure there’s got to be one for PDF generation.
- sznio 8mo ago>It should be much easier than that. You should should be able to serially test if each edit decodes to a sane PDF structure that's pointed out in the article. It's easy for plaintext sections, but not for compressed sections. Didn't notice any mention of checksums.
- chrisjj 8mo ago> it’s safe to say that Pam Bondi’s DoJ did not put its best and brightest on this Or worse. She did.
- eek2121 8mo agoI mean, the internet is finding all her mistakes for her. She is actually doing alright with this. Crowdsource everything, fix the mistakes. lol.
- rcakebread 8mo agoSicko.
- chrisjj 8mo agoLet's see her sued for leaking PII. Here in Europe, she'd be mincemeat.
- ISL 8mo agoThe US administration is, at present, regularly violating the law and ignoring court orders. Indeed, these very releases are patently in violation of multiple federal laws -- they're simultaneously insufficiently-responsive to meet the requirements of the law requiring the release of the files and fall afoul of CSAM laws by being incompletely redacted. The challenge, as we're all experiencing together, is that the law is not inherently self-enforcing.
- typeofhuman 8mo agoCan you provide a couple examples of the laws they're violating?
- mschuster91 8mo agoThere's more than enough credible reports of CSAM in the Epstein Files dump - more than enough for me to not go and download even a single file of them myself, simply because German law does not care about why you are in the possession of CSAM, even if you took the picture yourself. The legal situation regarding CSAM is very strict no matter which country, and I better hope no one here will actually be dumb enough to provide actual links.
- prettywoman 8mo ago[dead]
- pyrolistical 8mo agoIt decodes to binary pdf and there are only so many valid encodings. So this is how I would solve it. 1. Get an open source pdf decoder 2. Decode bytes up to first ambiguous char 3. See if next bits are valid with an 1, if not it’s an l 4. Might need to backtrack if both 1 and l were valid By being able to quickly try each char in the middle of the decoding process you cut out the start time. This makes it feasible to test all permutations automatically and linearly
- percentcer 8mo agoThis is one of those things that seems like a nerd snipe but would be more easily accomplished through brute forcing it. Just get 76 people to manually type out one page each, you'd be done before the blog post was written.
- WolfeReader 8mo agoYou think compelling 76 people to honestly and accurately transcribe files is something that's easy and quick to accomplish.
- altairprime 8mo agoNon-engineers are perfectly willing to volunteer their time to do drudgery. It's one of my opseng career's distinguishing specialties: I'll do drudgery rather than code when appropriate, rather than avoiding it or sulking about it (as was a common response at work for some number of decades!). Learned that lesson when I was 18 from an internship (where I completely failed to deliver any work product due to trying to code around the work). It's part of why I'm going into accounting: apparently having the stamina for dreary work is rare?! Also look up double/triple data-entry systems, where you have multiple people enter the data and then flag and resolve differences. Won't protect you from your staff banding together to fuck you over with maliciously bad data, but it's incredibly effective to ensure people were Actually Working Their Blocks under healthy circumstances.
- pbhjpbhj 8mo agoCaptcha!
- estimator7292 8mo agoFriend, have you ever heard of secretaries?
- fragmede 8mo ago> Just get 76 people I consider myself fairly normal in this regard, but I don't have 76 friends to ask to do this, so I don't know how I'd go about doing this. Post an ad on craigslist? Fiverr? Seems like a lot to manage.
- zahlman 8mo ago> …but good luck getting that to work once you get to the flate-compressed sections of the PDF. A dynamic programming type approach might still be helpful. One version or other of the character might produce invalid flate data while the other is valid, or might give an implausible result.
- deleted 8mo ago[deleted]
- yunnpp 8mo agoTime to flex those Leetcode skills.
- eek2121 8mo agoHonestly, this is something that should've been kept private, until each and every single one of the files is out in the open. Sure, mistakes are being made, but if you blast them onto the internet, they WILL eventually get fixed. Cool article, however.
- misja111 8mo agoWon't that entire DOJ archive already be downloaded for backup by several people? If I'd be a journalist working on those files, this is the very first thing I would do as soon as those files were published. Just to make sure you have the originals before DOJ can start adding more redactions.
- myduck_hacker 8mo ago[dead]
- bawolff 8mo agoTeseract supports being trained for specific fonts, that would probably be a good starting point https://pretius.com/blog/ocr-tesseract-training-data https://pretius.com/blog/ocr-tesseract-training-data
- kevin_thibedeau 8mo agopdftoppm and Ghostscript (invoked via Imagemagick) re-rasterize full pages to generate their output. That's why it was slow. Even worse with a Q16 build of Imagemagick. Better to extract the scanned page images directly with pdfimages or mutool. Followup: pdfimages is 13x faster than pdftoppm
- masfuerte 8mo agoThis. Not only is it faster, the images are likely to be of better quality. If you rasterize the pages then the images will be scaled, unless you get very lucky.
- velaia 8mo agoBummer that it's not December - the https://www.reddit.com/r/adventofcode/ https://www.reddit.com/r/adventofcode/ crows would love this puzzle
- blindriver 8mo agoOn one hand, the DOJ gets shit because it was taking too long to produce the documents, and then on another, they get shit because there are mistakes in the redacting because there are 3 million pages of documents.
- thereisnospork 8mo agoConsidering the justice to document ratio that's kind of on them regardless.
- rapind 8mo agoWhat they are redacting is pretty questionable though. Entire pages being suspiciously redacted with no explanation (which they are supposed to provide). This is just my opinion, but I think it's pretty hard to defend them as making an honest and best effort here. Remember they all lied about and changed their story on the Epstein "files" several times now (by all I mean Bondi, Patel, Bongino, and Trump). It's really really hard to give them the benefit of the doubt at this point.
- Rebelgecko 8mo agoMy favorite is that sometimes they redact the word "don't". Not only does it totally change the meaning of whatever sentence it's in, the conspiracy theory is that they had a Big Dumb Regex for redacting /Don\W+T/i to remove Trump references
- rexpop 8mo ago"On the one hand the chef gets shit for taking too long, and then on another for undercooked, badly plated dishes." Incompetence is incompetence.
- hypeatei 8mo agoThe zeitgeist around the files started with MAGA and their QAnon conspiracy. All the right wing podcasters were pushing a narrative that Trump was secretly working to expose and takedown a global child sex trafficking ring. Well, it turns out, unsurprisingly, that Trump was implicated too and that's when they started to do a 180. You can't have your cake and eat it too.
- legitster 8mo agoGiven how much of a hot mess PDFs are in general, it seems like it would behoove the government to just develop a new, actually safe format to standardize around for government releases and make it open source. Unlike every other PDF format that has been attempted, the federal government doesn't have to worry about adoption.
- deleted 8mo ago[deleted]
- derwiki 8mo agoJPEG?
- legitster 8mo agoThat's not really comparable - It needs to be editable and searchable.
- charcircuit 8mo agoPhotoshop and Google Images show it can be done.
- recursive 8mo agoLossy
- iberator 8mo agoPNG
- Spooky23 8mo agoYou’re thinking about this as a nerd. It’s not a tools problem, it’s a problem of malicious compliance and contempt for the law.
- legitster 8mo ago
- nubg 8mo agoWait would this give us the unredacted PDFs?
- poyu 8mo agoI think it's the PDF files that were attached to the emails, since they're base64 encoded.
- ryanSrich 8mo agoThat's the idea yeah. There are other people actively working on this. You can follow vx-underground on twitter. They're tracking it.
- sznio 8mo agoFrom the unredacted attachments you could figure out what the redacted content most likely contains. Just like the other sloppy redactions that sometimes hide one party of the conversation, sometimes the other, so you can easily figure out the both sides.
- deleted 8mo ago[deleted]
- dperfect 8mo agoNerdsnipe confirmed :) Claude Opus came up with this script: https://pastebin.com/ntE50PkZ https://pastebin.com/ntE50PkZ It produces a somewhat-readable PDF (first page at least) with this text output: https://pastebin.com/SADsJZHd https://pastebin.com/SADsJZHd (I used the cleaned output at https://pastebin.com/UXRAJdKJ https://pastebin.com/UXRAJdKJ mentioned in a comment by Joe on the blog page)
- pests 8mo agoSo it was a public event attended by 450 people: https://www.mountsinai.org/about/newsroom/2012/dubin-breast-center-holds-inaugural-gala https://www.mountsinai.org/about/newsroom/2012/dubin-breast-... https://www.businessinsider.com/dubin-breast-center-benefit-2012-12 https://www.businessinsider.com/dubin-breast-center-benefit-... Even names match up, but oddly the date is different.
- elmomle 8mo agoYour links are for the inaugural (first) ball in December 2011; OP's text referred to a second annual ball in December 2012.
- pests 8mo agoYou are right my first is incorrect but the second does seem to be from 2012.
- sorbus-25 8mo agoDUBIN BREAST CENTER SECOND ANNUAL BENEFIT MONDAY, DECEMBER 10, 2012 HONORING ELISA PORT, MD, FACS AND THE RUTTENBERG FAMILY HOST CYNTHIA MCFADDEN SPECIAL MUSICAL PERFORMANCES CAROLINE JONES, K'NAAN, HALEY REINHART, THALIA, EMILY WARREN MANDARIN ORIENTAL 7:00PM COCKTAILS LOBBY LOUNGE 8:00PM DINNER AND ENTERTAINMENT MANDARIN BALLROOM FESTIVE ATTIRE
- Groxx 8mo ago
- ChocMontePy 8mo agoYou can use the justice.gov search box to find several different copies of that same email. The copy linked in the post: https://www.justice.gov/epstein/files/DataSet%209/EFTA00400459.pdf https://www.justice.gov/epstein/files/DataSet%209/EFTA004004... Three more copies: https://www.justice.gov/epstein/files/DataSet%2010/EFTA02153691.pdf https://www.justice.gov/epstein/files/DataSet%2010/EFTA02153... https://www.justice.gov/epstein/files/DataSet%2010/EFTA02154109.pdf https://www.justice.gov/epstein/files/DataSet%2010/EFTA02154... https://www.justice.gov/epstein/files/DataSet%2010/EFTA02154246.pdf https://www.justice.gov/epstein/files/DataSet%2010/EFTA02154... Perhaps having several different versions might make it easier.
- ChocMontePy 8mo agoAlso, I found a different base64 encoding with a different font here: https://www.justice.gov/epstein/files/DataSet%209/EFTA00775520.pdf https://www.justice.gov/epstein/files/DataSet%209/EFTA007755... This doesn't solve the "1 & l" problem for the pdf you are looking at, but it could be useful anyway.
- ChocMontePy 8mo agoAnd this might be a copy of the original pdf: https://www.justice.gov/epstein/files/DataSet%2011/EFTA02702727.pdf https://www.justice.gov/epstein/files/DataSet%2011/EFTA02702...
- Aloisius 8mo agoI checked and that's definitely the black and white version of the one encoded in the file. Someone build a very simple OCR tool that successfully extracted the base64[1]. The only difference besides the lack of the document tracking ids at the bottom is the original was pink on blue for the first page and has some pink text on the second. [1] https://github.com/KoKuToru/extract_attachment_EFTA00400459 https://github.com/KoKuToru/extract_attachment_EFTA00400459
- JKCalhoun 8mo agoFile is gone now, hmmm…
- bushbaba 8mo agoThis proves my paranoia that you should print and rescan redactions. That or do screenshots of the pdf redacted and convert back to a pdf
- darig 8mo ago[dead]
- Snoozus 8mo agothis would not have helped here
- phanimahesh 8mo agoHow would that help in this case?
- Evidlo 8mo agoI took at stab at training Tesseract and holy jeebus is their CLI awful. Just an insanely complicated configuration procedure.
- subscribed 8mo agoGods, I had a flashback just from you mentioning that. I had a reasonably simple problem to solve, slightly weird font and some 10 words in English (I actually only missed one or two blocks for missing letters to cover all I needed). After a couple of days having almost everything (?) I just surrendered. This seems to be intentionally hostile. All the docs scattered across several repositories, no comprehensive examples, etc. Absolutely awful piece of software from this end (training the last gen).
- queenkjuul 8mo agoI'm only here to shout out fish shell, a shell finally designed for the modern world of the 90s
- deleted 8mo ago[deleted]
- sorbus-25 8mo agoEvent details: https://web.archive.org/web/20260206040716/https://what2wearwhere.com/dubin-breast-center-2nd-annual-benefit/ https://web.archive.org/web/20260206040716/https://what2wear...
- sorbus-25 8mo agoDUBIN BREAST CENTER SECOND ANNUAL BENEFIT MONDAY, DECEMBER 10, 2012 HONORING ELISA PORT, MD, FACS AND THE RUTTENBERG FAMILY HOST CYNTHIA MCFADDEN SPECIAL MUSICAL PERFORMANCES CAROLINE JONES, K'NAAN, HALEY REINHART, THALIA, EMILY WARREN MANDARIN ORIENTAL 7:00PM COCKTAILS LOBBY LOUNGE 8:00PM DINNER AND ENTERTAINMENT MANDARIN BALLROOM FESTIVE ATTIRE
- sorbus-25 8mo agoSome pics from the event. Doppelgänger in the background?: https://web.archive.org/web/20121215131412/https://thaliadiva.wordpress.com/2012/12/11/nuevas-fotos-thalia-en-the-dubin-breast-center-2nd-annual-benefit-10-12-2012/ https://web.archive.org/web/20121215131412/https://thaliadiv...
- SomaticPirate 8mo agoAre there archives of this? I have no doubt after this post goes viral some of these files might go “missing” Having a large number of conspiracies validated has lead me to firmly plant my aluminum hat
- direwolf20 8mo agohttps://github.com/yung-megafone/Epstein-Files https://github.com/yung-megafone/Epstein-Files
- winddude 8mo agohere's another few to decode, https://www.justice.gov/epstein/files/DataSet%2010/EFTA01804740.pdf https://www.justice.gov/epstein/files/DataSet%2010/EFTA01804... https://www.justice.gov/epstein/files/DataSet%209/EFTA00775520.pdf https://www.justice.gov/epstein/files/DataSet%209/EFTA007755... https://www.justice.gov/epstein/files/DataSet%209/EFTA00434905.pdf https://www.justice.gov/epstein/files/DataSet%209/EFTA004349... and than this one judging by the name of the file (hanna something) and content of the email: "Here is my girl, sweet sparkling Hanna=E2=80=A6! I am sure she is on Skype " maybe more sinister (so be careful, i have no ideas what the laws are if you uncover you know what trump and Epstein were into)... https://www.justice.gov/epstein/files/DataSet%2011/EFTA02715081.pdf https://www.justice.gov/epstein/files/DataSet%2011/EFTA02715... [Above is probably a legit modeling CV for HANNA BOUVENG, based on, https://www.justice.gov/epstein/files/DataSet%209/EFTA01120466.pdf https://www.justice.gov/epstein/files/DataSet%209/EFTA011204..., but still creepy, and doesn't seem like there's evidence of her being a victim]
- Snoozus 8mo agothis one has a better font, might be a simple copy&paste job
- winddude 8mo agoI've checked for copy and paste, there's so many character flaws, their OCR must have sucked really bad, I may try with deepseekOCR or something. I mean the database would probably more searchable if someone ran every file through a better OCR.
- netsharc 8mo agoGeezus, with the short CV in your profile, you couldn't tell an LLM to decode "filename=utf-8"CV%5F%5F%5FHanna%5FTr%C3%A4ff%5F.pdf"? That's not "Bouveng". Anyway searching for the email sender's name, there's a screenshot of an email of hers in English offering him a girl as an assistant who is "in top physical shape" (probably not this Hanna girl). That's fucking creepy: https://www.expressen.se/nyheter/varlden/epsteins-lofte-till-barbro-ehnbom-for-att-fa-en-kvinnlig-assistent/ https://www.expressen.se/nyheter/varlden/epsteins-lofte-till...
- ks2048 8mo agoI wonder if jmail (https://www.jmail.world/ https://www.jmail.world/) has worked on this? I tried to find the message in this blog post, but couldn't. (don't see how to search by date).
- heraldgeezer 8mo ago[flagged]
- nullorempty 8mo agoit's really all about the blackmail
- wtcactus 8mo agoMy non political take about this gift that keeps on giving is that: PDF might seem great for the end user that is just expected to read or print the file they are given, but the technology actually sucks. PDF is basically a prettify layer on top of the older PS that brings an all lot of baggage. The moment you start trying to do what should be simple stuff like editing lines, merging pages, change resolution of the images, it starts giving you a lot of headaches. I used to have a few scripts around to fight some of its quirks from when I was writing my thesis and had to work daily with it. But well, it was still an improvement over Word.
- direwolf20 8mo agoIt's meant as a printer replacement format, hence "print to PDF". It's a computer file format about equivalent to a printed document. Like a printed document, you can't just change its structure and recompile it.
- deleted 8mo ago[deleted]
- IshKebab 8mo agoDisappointing how terrible open source OCR still is.
- tcgv 8mo ago> Then my mom wrote the following: “be careful not to get sucked up in the slime-machine going on here! Since you don’t care that much about money, they can’t buy you at least.” I'm lucky to have parents with strong values. My whole life they've given me advice, on the small stuff and the big decisions. I didn't always want to hear it when I was younger, but now in my late thirties, I'm really glad they kept sharing it. In hidhsight I can see the life-experience / wisdom in it, and how it's helped and shaped me.
- pavel_lishin 8mo agoI think this was meant to be a reply to https://news.ycombinator.com/item?id=46903929 https://news.ycombinator.com/item?id=46903929 ?
- tcgv 8mo agoIndeed! Thanks for pointing that out. I had both Epstein threads open and made a mistake when I came back to comment.
- alhamdulillah23 8mo agoGot it. Page 1: https://imgur.com/a/jwgu9uH https://imgur.com/a/jwgu9uH Page 2: https://imgur.com/a/4Zi3bkk https://imgur.com/a/4Zi3bkk Use this: https://github.com/KoKuToru/extract_attachment_EFTA00400459 https://github.com/KoKuToru/extract_attachment_EFTA00400459