5 ms·
very interesting. the windows 10 screen reader consume the raster data on PDFs to OCR and not the code point data embedded within the PDF. People here have been
by viggity 4y ago
very interesting. the windows 10 screen reader consume the raster data on PDFs to OCR and not the code point data embedded within the PDF. People here have been on my ass about saying "OCR resistant" and I get where they are coming from. I've primarily been testing the various "OCR" functionalities built within the various PDF readers out there. The "OCR" that 98% of laypeople are going to rely on. I always new that exporting to an image based PDF wouldn't be defeated. If a human can read it, a machine can read it. Just most PDF readers aren't set up to do it. Out of curiosity, when you use your screen reader on my website, does the <textarea> read and/or start with "Name: Satoshi Nakamoto"?
- ctoth 4y agoI can see the content in the textareas are a bunch of Unicode glyphs that aren't mapped to speakable characters and when I perform a "read all" action mostly render as questionmarks.
- viggity 4y agoInteresting how the screen reader will use the raster image for a PDF but not for the website. Based on the work I had to do to get the PDF to work, PDFs are a clusterfuck of glyphs floating in space and not really structured like HTML, so presumably they have to do the raster for the PDF so they can determine+optimize their own text flow but they're relying on the hierarchy of HTML to be enough of a guidepost that they don't need to OCR raster data. Thank you very much for replying!
- deleted 4y ago[deleted]
- solardev 4y agoI don't think people have an issue with your implementation, but your misrepresentation. OCR has a specific meaning and this is not resistant to it (at all), in fact it encourages people to do it because the text isn't already copyable. What your service DOES block is casual copying and scraping. But the people who are going to be doing that (search engines and the like) are different from people who need actual protection from OCR (I don't know who that is, but presumably they've identified it as a threat and need to specifically mitigate it, a la CAPTCHAs). By misrepresenting the actual security/obscurity of your service, you are putting people at risk with a false sense of security that's trivially defeated by anybody with minimal IT experience. It'd be like if Signal promised encryption but actually just implemented ROT13. Which would you rather hear from, a bunch of devs saying "you're misrepresenting your product, might wanna tweak your marketing" or a bunch of burned users trying to sue you because you misled them into a bad situation?
- weird-eye-issue 4y ago"If a human can read it, a machine can read it." Completely contradicts the clickbait title I wish I never clicked