12 ms·
Show HN: PDFs that are readable by human eyes only
Hi, OP here. A friend was involved in a custody battle and was afraid his ex was going to leak all of his discovery documents on the internet and he asked if there was something I could do to make it harder for bots/crawlers to find sensitive information. Originally I was going to turn all of his docs to image based PDFs, but those get large fast and are easy to OCR.
So I found a post musing about altering fonts/glyphs so that it looks like english, but the actual character being seen by the pdf reader is a non-english character. As such, when you try to OCR these files, it doesn't see any images and can't convert it.
I figured it had some potential uses and maybe you fine folks can identify other use cases. I'll be monitoring this post most of the day.
- deleted 4y ago[deleted]
- maybeiambatman 4y agoFascinating. How does this work?
- viggity 4y agoIt takes the embedded font out of your PDF, and then maps non-latin characters (japanese, cyrillic, etc) to render as if they looked like a latin character. So in the example on the site. "ӕ" will render as a "D" using my special font. And "ㅈ" will draw the "B" glyph. Then I do a replacement on the underlying text so all "B" are replaced with "ㅈ". It is more complicated than that, but that's the gist.
- bambax 4y agoSo basically it's a type of Caesar cipher where letters are mapped to something else one-to-one. Very easy to decrypt / reverse. If this tool ever became popular there would be hundreds of scripts to defeat it. And as it is, it does not prevent "OCR", only copy-paste.
- smegsicle 4y agolooks like it's actually one-to-many across unicode, if so then you could think of it as approaching one-time-pad encryption, with the key being the font if the generator crafted a new font every time, never used the same codepoint twice, and kept the font separate from the document (pre-shared by being installed on the intended receiver's machine) then it'd be uncrackable!
- blitz_skull 4y agoI literally used an OCR tool to grab the text directly out of the first box. I think this is meant to be guarding against copy/pasting—not OCR.
- viggity 4y agoSo this is interesting. I guess I didn't realize that there are (common?) tools to OCR screenshots. And do that end, there probably isn't a whole lot I can do to stop it. But when you're looking at a huge tax return, or sworn testimony, or just a dump of 3000 emails, you're not gonna screenshot each one. You're going to want to automate the OCR, which most PDF readers (at least the commercial ones) will let you do. It is against that type of OCR that my app is resistant to. They look for image data within the PDF and OCR that. They bypass my text because to the pdf reader, it already is in a text format. I'm 1000% sure there are gurus who could whip up a script to overcome this. But its kind of one of those things where you don't have to outrun the bear, you have to outrun your friend running next to you. It makes your sensitive documents just that much less likely to be scanned/found.
- bil7 4y ago> I didn't realize that there are (common?) tools to OCR screenshots. This seems like quite the oversight to me...
- viggity 4y agoBut if you have a huge tax document, you're likely not going to screenshot page by page. Yes, there are ways to automate this. But if you're 50 year old divorce attorney, you're going to click on the "OCR" button in your PDF reader and it will not work.
- Semaphor 4y agoI’m not sure what reader you are talking about, but that button is most certainly not doing any kind of OCR if your technique stops it.
- solardev 4y agoHmm... interesting in theory, but take a screenshot and it's trivially bypassed. Try it yourself here: http://www.structurise.com/screenshot-ocr/ http://www.structurise.com/screenshot-ocr/
- viggity 4y agoI posted some more info here: https://news.ycombinator.com/item?id=32003066 https://news.ycombinator.com/item?id=32003066
- rst 4y agoWhich is not what I'd expect from anything that claims to be "OCR resistant". It's not at all clear what they mean by that.
- solardev 4y agoI think the OP, while well-intentioned, did not really understand how OCR works. Follow-up convo in a separate thread here: https://news.ycombinator.com/item?id=32003066 https://news.ycombinator.com/item?id=32003066 What this blocks is not OCR but casual copy & pasting (and search engine indexing)
- RajT88 4y agoI think it works for the use case - where documents can be provided for discovery, but if posted online won't have the content indexed by search engines. The various legal teams involved are unlikely to ever be the wiser. Or will they? Won't this print out a pile of gibberish? Hard copies are rather important in the courts. Somebody is going to complain about what was provided in that case.
- londons_explore 4y ago> his ex was going to leak all of his discovery documents on the internet If that happens, I suspect he'd have a very strong case to win custody...
- viggity 4y agohe got what he wanted out of the case, so good for him. The problem is that he just didn't know if it was going to happen. And she could definitely get sanctioned for it. But his info would still be out there.
- mbreese 4y agoTo me, this falls under the category of — you can’t have a technical solution to a societal problem. Yes, technology may have made your friend feel better. But the actual thing that protected him was the law, not the obfuscated PDFs. But, if the judge didn’t care and it made your friend feel better, who am I to judge? But this isn’t a great protection scheme… it just adds a few extra technical hurdles that are easy to get around.
- waynesonfire 4y agobrilliant solution.
- solardev 4y ago> As such, when you try to OCR these files, it doesn't see any images and can't convert it. That isn't true. Acrobat might skip parts of the PDF that it thinks are already text/glyphs, but it's trivial to get around that by either using other OCR software or just printing the PDF to a raster image first. Example: https://filebin.net/qse2e0oaqkl1hjof/ocred.pdf https://filebin.net/qse2e0oaqkl1hjof/ocred.pdf Still, though, for the purposes of obscuring these from bots/crawlers... a lot better than nothing!
- anyfactor 4y agoI am on a older phone. What your eyes see is identical to what the computer sees. They are both giberrish. Also the email is giberrish. What am I missing here?
- viggity 4y agoI'm wondering if your browser isn't showing woff fonts for some reason. The "your eyes see" textbox on a modern browser shows: Name: Satoshi Nakamoto DOB: 1982-06-05 SSN: 958-20-3141 Cell Phone: 514-867-5309
- solardev 4y agoYour PDF reader probably doesn't support the particular font this is using. Try it on your computer.
- josephcsible 4y agoThis is basically just really weak DRM and is just as evil.
- maxbond 4y agoI agree that OCR is an important tool for end users, especially those with accessibility needs, and that we shouldn't use something like this lightly, but the context is completely different here. If I want to send you a PDF from my gmail, and want to make it difficult for Google to leverage that data - that's completely different than if I were a giant media company, gatekeeping to ensure a huge portion of culture flows through me, which I then claim as my own and charge exorbitant rents for, enforced by DRM. The problem with DRM is not that someone is trying to control what happens to a string of bits, it's that it props up an institution which is harmful.
- withinboredom 4y agoThe problem is that people won’t “use it lightly” and not everyone speaks the same language. Being able to copy/paste text into a translation tool (probably Google Translate which is kinda ironic in the case you mentioned) to understand what the document is about is super important when in another country or communicating with someone in another country.
- maxbond 4y agoBeing able to copy/paste is an important option that empowers users. Being able to selectively defeat copy paste is an additional option that additionally empowers users. I don't anticipate this tool being used very widely. If it became the default, I would have a problem with it for the reasons you highlight, among others.
- josephcsible 4y agoIt doesn't empower users when the final authority over what a computer does or refuses to do lies with anyone other than the computer's owner.
- O__________O 4y agoYou asked about use case ideas, while I personally strongly dislike them, there are number of sites, including online testing apps, that try to remove copy-and-paste. Not sure how valuable it would be, but it’s for sure a use case.
- thrown_22 4y agoCybersecurity Incident & Vulnerability Response Playbooks Operational Procedures for Planning and Conducting Cybersecurity Incident and Vulnerability Response Activities in FCEB Information Systems Publication: November 2021 Cybersecurity and Infrastructure Security Agency DISCLAIMER: This document is marked TLP:WHITE. Disclosure is not limited. Sources may use TLP:WHITE when information carries minimal or no foreseeable risk of misuse, in accordance with applicable rules and procedures for public release. Subject to standard copyrght rules, TLP:WHITE information may be distributed without restriction. For more information on the Traffic Light Protocol, see --- Converting the first page of the sample PDF file to a tiff file using ghost script and running tesseract OCR without any special filters. >Resistant to Optical Character Recognition (OCR), most laypeople will need to print+rescan to OCR This is not OCR resistant, I used the same two liner I used to get my textbooks scanned at university 20 years ago.
- viggity 4y agoYou specifically are technologically proficient. Not everybody knows how to export to a tiff and then OCR. 98% don't. When I say "OCR Resistant", I mean that I haven't found PDF software with built in OCR that has managed to extract the english text back out.
- thrown_22 4y agoThat's like saying a lock is pick resistant because you haven't been able to open it with a dead fish. Words mean things, if what you did can't stand up to 20 year old technology then it's basically useless. Remove the claim that it resists OCR and just called it copy/paste proof and unsearchable.
- smegsicle 4y agoif a lock convinces most popular lock-picking devices to use the ineffective dead fish technique then it's something atleast
- 4y ago
- oxff 4y agoJust give me readable papers instead. Such a pain in the dick format yet its the only thing there is.
- dodo6502 4y agoAs an author of a PDF library this is hilarious, because the number of bugs I have received over the years where this is unintentionally happening is quite high.
- revolvingocelot 4y agoIf someone who actually has to deal with the PDF standard can't help OP, I don't think anybody can.
- jffry 4y agoI think people will get tripped up by you saying it "can't" be OCR'd or that it is difficult to do so, and will end up looking past a pretty elegant solution in the process. This seems like a nicely clever way to trip up non-targeted scrapers which might attempt to OCR any images they encounter, but which will ignore what looks like random gibberish codepoints. It doesn't eliminate the ability to index this data but I can see how it might greatly reduce it. Obviously you could still convert these PDFs to an image and OCR them, but that's not the thing being defended against here.
- forgotpwd16 4y agoNot sure OCR is the correct term here. OCR specifically means extracting text from an image. This approach doesn't protect against that. Some maybe better options will be "machine obfuscated" or "scrape resistant".
- peanut_worm 4y agoThis seems like it would just be annoying and would not even work for most purposes. Kind of neat though.
- ksaj 4y agoI'm not very convinced by the PDF idea, but web fonts done this way would be great for the parts of web pages you don't want scraped or collected by search engines, if it is on pages where you do want at least some of the content available to search engines.
- viggity 4y agothe antiscraping thing is a good idea. Hell, you could poison it very lightly using homoglyphs (greek capital Epsilon instead of just an E) just to see where else on the internet your data ends up, too.
- mmastrac 4y agoFunny enough, this is because the PDF spec literally allows you to map glyphs like that. Some properly-produced PDFs are broken like this, but it's been less common in recent years. You're supposed to provide mapping tables for text extraction but they are optional. This fails pretty bad for security because you can detect the glyphs themselves in the font tables and provide a mapping yourself
- nemothekid 4y agoI thought this was going to be some adversarial neural network. Unfortunately OP, I don't think your solution even works for your intended use case; Google already does OCR (actual OCR, not just parsing text) in Images. I use it in Gmail quite often. Regardless the implementation is quite neat and will surely thwart less advanced indexers.
- social_quotient 4y agoDoes it ocr on live text pdfs or just pdfs that have text in images/flattened?
- ben_w 4y agoLikewise iOS, text in screenshots is selectable and in this case it is recognised correctly.
- nmstoker 4y agoAdding to this, it's trivial to get the human readable text on a Google phone: switch between apps and pick the app with the text but don't jump back into the app yet, select text and you can immediately copy the text out. It's yielding the visible text, so most be OCR'ing the image (works offline too). When i copy direct from the example on the web page, in the non-OCR method, that does give the messed up text, but not when done the way above. " Phone: 514-867-5309" was copied out easily (can't be bothered to go back get the Cell bit i was just inaccurate copying, I'm sure it works!)
- viggity 4y agothank you to everyone in this thread who realize it was never meant to be perfect and appreciating it for what it does do!
- SilasX 4y agoThe point isn't that it has flaws, but that its description is wrong. "Non-human eyes" -- normally understood to be OCR -- read it just fine. I think most of us were expecting something that disrupts "computer eyes" (e.g. because of deceiving overly narrow "tricks" that neural networks use to identify characters) but left it readable for the typical human (like an easy Captcha). A more accurate (and helpful!) description of the problem you're solving is that this disrupts text parsers. That is, any program that just reads this in as text won't see the "real" letters (unless it's been pre-programmed with a specific reverser, etc.) and thus will frustrate, say, text search. Which, on that note, I notice elsewhere you mention this being a solution applied to document submission in legal proceedings. In that case, the assumption might be that one side wishes to run text searches and assume its compatible with that. In that case, this could be viewed as non-compliance with a judge's orders, so FYI.
- Komodai 4y ago"As such, when you try to OCR these files, it doesn't see any images and can't convert it." Bullshit 1. Screenshot 2. OCR 3. Profit
- hexo 4y agoImpressive, the example web text area is very lovely too
- dawnerd 4y agoThis is terrible for accessibility, please don't do this.
- vehemenz 4y agoIsn't that the point though? If you make a sensitive document unindexable (assuming this works), then effectively no one can find it. The intent here is not to restrict the document to sighted users but to hide the document from everyone, which includes sighted and blind users searching for keywords. The fact that blind users can't read the text at all without screen grabbing is just a bonus.
- dawnerd 4y agoTreating blind people that way will 100% lead you towards a lawsuit. So have fun with that I guess.
- hk1337 4y agoI thought this was going to be some style guide on how to make the PDF document easy on the eyes to read.
- layer8 4y agoYou could use the ZXX typeface to defy OCR: https://walkerart.org/magazine/sang-mun-defiant-typeface-nsa-privacy https://walkerart.org/magazine/sang-mun-defiant-typeface-nsa... It’s probably still not ML-resistant.
- rafram 4y agoYou could just train an OCR engine on that typeface. IIRC training Tesseract for a new font is quite trivial.
- jeroenhd 4y agoOnce you get it working, I'm sure. I've tried doing exactly that and every time I end up scouring through unmaintained scripts, old manuals and help guides that assume you're already intimately familiar with the tool itself. In the end I managed to get some text out of it but the end result was still pretty terrible.
- ctoth 4y agoI am a blind user using an extension to my screen reader which (under the covers) uses the Windows 10 built-in OCR. Your sample document gives me: INTRODUCTION The Cybersecurity and Infrastructure Security Agency (CISA) is committed to leading the response to cybersecurity incidents and vulnerabilities to safeguard the nation's critical assets. Section 6 of Executive Order 14028 directed DHS, via CISA, to "develop a standard set of operational procedures (playbook) to be used in planning and conducting cybersecurity vulnerability and incident response activity respecting Federal Civilian Executive Branch (FCEB) Information Systems." I Overview This document presents two playbooks: one for incident response and one for vulnerability response. These playbooks provide FCEB agencies with a standard set of procedures to identify, coordinate, remediate, recover, and track successful mitigations from incidents and vulnerabilities affecting FCEB systems, data, and networks. In addition, future iterations of these playbooks may be useful for organizations outside of the FCEB to standardize incident response practices. Working together across all federal government organizations has proven to be an effective model for addressing vulnerabilities and incidents. Building on lessons learned from previous incidents and incorporating industry best practices, CISA intends for these playbooks to evolve the federal government's practices for cybersecurity response through standardizing shared practices that bring together the best people and processes to drive coordinated actions. Pretty sure this doesn't actually work.
- viggity 4y agovery interesting. the windows 10 screen reader consume the raster data on PDFs to OCR and not the code point data embedded within the PDF. People here have been on my ass about saying "OCR resistant" and I get where they are coming from. I've primarily been testing the various "OCR" functionalities built within the various PDF readers out there. The "OCR" that 98% of laypeople are going to rely on. I always new that exporting to an image based PDF wouldn't be defeated. If a human can read it, a machine can read it. Just most PDF readers aren't set up to do it. Out of curiosity, when you use your screen reader on my website, does the <textarea> read and/or start with "Name: Satoshi Nakamoto"?
- ctoth 4y agoI can see the content in the textareas are a bunch of Unicode glyphs that aren't mapped to speakable characters and when I perform a "read all" action mostly render as questionmarks.
- peetah 4y agofunny, I did this for the web, something like 8 or 7 years ago, under the name "cprotext", but was unable to find a way to sell this as SaaS :) There should be a wordpress plugin floating around somewhere called wp-cprotext and maybe one or two demo websites that I can't even remember the url. People came with the same critics as we can read here: evil DRM, accessibility nightmare and easily bypassed by OCR. All in all, I came to be quite convinced by these critics, especially the first and second, and shut it down completely. I would genuinely be interested to see how you'll succeed where I failed ! good luck !
- viggity 4y agoInteresting. Thanks for the info. I'll look it up!
- bscphil 4y agoThis is broken in multiple ways, some obvious, some not. 1. Obviously most people jumped directly to OCR, and that works of course. So counter to the OP, you can trivially render the first page to a high resolution PNG and then OCR that with what will probably be 100% accurate results. Sample image: https://i.imgur.com/hyJOSjY.jpg https://i.imgur.com/hyJOSjY.jpg 2. This is just messing with glyphs in fonts, so one trivial way of undoing the changes losslessly (not even requiring OCR) is to create a mapping between each font glyph and the original character. For example, I was able to extract the font used for the text "CONTENTS" near the beginning of the sample document. It is named "SecureFont-1845559949-FranklinGothic-Demi", as extracted by mutool. In the PDF "CONTENTS" is made up of eight Unicode characters, which render as "CONTENTS" in that font. 3. Even if the first two methods somehow failed, the same character in a given font is repeatedly used to render the same character in English. That makes the approach similar to a substitution cipher [1] which is trivially broken with frequency analysis. You could literally just copy / paste the fake "text" out of the PDF and with an analysis tool derive the original text. This isn't really significant since the PDF can be read by sight anyway, but it's worth pointing out. [1] https://en.wikipedia.org/wiki/Substitution_cipher https://en.wikipedia.org/wiki/Substitution_cipher
- rmbyrro 4y agoI believe this can be taken as an axiom: there's nothing a human eye can see that a machine cannot.
- mdaniel 4y agoIsn't the opposite of your assertion the whole reason reCaptcha (and its dumbass hcaptcha competitor) exists? Maybe you were asserting only on written text, and not machine vision in the general case -- but even then I'd bet that only applies to very regular text, and not handwritten items
- hunterb123 4y agoAre you asserting that reCaptcha is not solvable by OCR? It's a two parter if so. First, the main reCaptcha checks are via JS, that's what determines your likelihood to be a bot and the difficulty of the check. That's a cat and mouse game, but ultimately the client wins. Second is the images like road signs, red lights, etc. can be solved via OCR, and there are services that do so. But reCaptcha keeps a decent amount of the opportunistic actors out, so it's good enough for the industry. But I will give you that currently the human captcha services have a better solve rate, albeit they are more expensive, but it's been getting closer and closer and it's more economical to use an OCR service.
- s1mon 4y agoMacOS does OCR on this just fine. Screenshot, open in Preview, and select, copy, paste: Name: Satoshi Nakamoto DOB: 1982-06-05 SSN: 958-20-3141 Cell Phone: 514-867-5309
- donkarma 4y agotook me a while to realise that my font settings on firefox break this
- loxias 4y agoI like your intent (helping your friend), and I'm _REALLY_ impressed at the polish (clever name, nice looking website, convert-for-free microservice) but I'm sad to report (like other comments here) that in practice this is useless. :/ Repeating what others have said, don't assume that anyone who cares will even bother looking at the PDF's embedded text. They'll rasterize then OCR. To scale up, they'll just deploy more tesseract pods. :) (At least, this is how I've seen it done!) I'd take what you have, tweak/rebrand a bit. I personally don't use embedded text in PDFs for anything, I know firsthand some large crawlers don't either, but perhaps something does. Identify that one $foo that uses embedded text then rebrand as a "$foo obfuscator" :) Alternatively, you could try to make something that really does confuse OCR! You can't foil everyone but you could raise the barrier to entry. Most rely on the pre-trained model, which you can also use. Keep permuting the image until the resulting PDF gives garbage when you try to OCR it. I'm sure you can do all sorts of transforms to the PDF that make the resulting image ugly-but-readable to humans, but really frustrating to use off-the-shelf OCR on. Mess with spacing then add slightly colored light geometric shapes in the white space. Change image contrast slowly from one corner to another. Things like that. ;)
- solardev 4y agoHaving to read 40 pages of captchas is sure to please the judge :)
- Pakdef 4y agoWith Javascript disabled, it doesn't work... they both look about the same
- hartator 4y agoUsing CLI OCR, I am getting a 100% match: `Name: Satoshi Nakamoto DOB: 1982-06-05 SSN: 958-20-3141 Cell Phone: 514-867-5309`
- schoen 4y ago> afraid his ex was going to leak all of his discovery documents on the internet Any chance of getting a clear protective order in advance from the judge, and then enforcing that with contempt of court sanctions if his ex violated the order? The discovery process is very, very intrusive (sometimes unbelievably so), but judges presiding over cases involving sensitive discovery may be interested in trying to use their powers to mitigate harms from leaks. (You could use some kind of PDF watermark or text watermark to confirm that the leaked version was the version produced in discovery.)
- deleted 4y ago[deleted]
- oa335 4y agoI like the idea, it’s a decent approach to stop low-effort scraping of your info.
- hda2 4y agoPassword-protect (i.e. encrypt) the documents and provide the passwords on a separate piece of paper with your reasoning and concerns clearly laid out. This won't prevent the leaking of those documents, but it will prevent automatic indexing by search engines unless they deliberately strip out encryption prior to leaking them. If they do strip it out, then the excuses of them "accidentally" leaking the documents becomes very implausible to any reasonable judge. For good measure, add unique watermarks to all documents to make it easier to prove who leaked the documents later.
- cordite 4y agoMy AP history textbooks did this, yet somehow was still searchable with ctrl-f. I wonder what the compromise was.
- deleted 4y ago[deleted]
- ehhthing 4y agoLesser known fact: this "feature" is actually built into Chrome on MacOS! If you try to print a website as PDF it will completely break it and copy and paste will result in random Unicode characters.
- maxloh 4y agoBut why did this "feature" added in the first place? In most cases, this feature isn't needed but causes confusion instead.
- bryanrasmussen 4y agoAside from it evidently not working there is, it seems to me, a logical problem at the root - if your friend's x was going to leak all the documents on the internet and it was readable by people and not crawlers people would still read it, probably transcribe it if your friend was important enough to worry about this kind of thing, and then put the transcribed documents online. It seems more likely that if you wanted to pass around documents with hidden content you would do so via the time honored methods of steganography and not anything else.
- secretsatan 4y agoI thought it would be trivial to render and apply OCR and was going to give it a quick attempt in Swift, but looking at the comments, looks like many implementations already exist without even trying
- justshowpost 4y agoThe reason - to make it only readable by human eyes - doesn’t make sense. It is always only ever readable by human eyes since bots don’t have eyes. You don’t mean eyes, but consciousness? Bots don’t have consciousness, at least until AGI isn’t here yet. By being indexed by bots and published on the Web, it’s just more human eyes. If you just want certain human eyes, then you instead need access control. By trying to exclude machines you just make it harder for humans to use the information, like copy pasting, text-to-speech, screen readers for the blind, etc.
- hedora 4y agoHas anyone tried uploading one of these pdfs to a website? I’m curious to know whether search engines successfully ocr and index stuff like this already.
- _8j50 4y agoPhishers and scammers will like this. If they don't have links they resort to using pictures of text instead and hope for lack of OCR scanning. But I am sure this can be also defeated by OCR.