6 ms·
> For the blind, PDF is the worst possible format I'm surprised. Nobody has created an accessibility solution for PDFs after all these years of ubiquity? What'
by hackuser 10y ago
> For the blind, PDF is the worst possible format
I'm surprised. Nobody has created an accessibility solution for PDFs after all these years of ubiquity? What's the story?
- cygx 10y agoThey probably can get at the text, but I have doubts about any formulae...
- robin-berjon 10y agoPDF is only accessible if it is specifically crafted to be so (and then again, I don't think that goes very far). To the best of my knowledge only Adobe's tools and Word actually output that (and at least in the latter they tend to look like train-wrecks, which may or may not be PDF's fault more than Word's).
- contravariant 10y agoWell, I don't know about PDF, but PostScript, which it's based on is basically just a programming language for drawing symbols in specific positions on a page. Depending on how it was written (or more likely, generated) this could be readable, or incredibly unreadable. As an example, in post script the following example from Wikipedia would simply show the text Hello World: %!PS /Courier % name the desired font 20 selectfont % choose the size in points and establish % the font as the current one 72 500 moveto % position the current point at % coordinates 72, 500 (the origin is at the % lower-left corner of the page) (Hello world!) show % stroke the text in parentheses showpage % print all on the page Now if you're lucky you can just extract all quoted text and read those in order, but that's unlikely to work for all documents.
- hackuser 10y agoAlmost all formats would similarly require content to be parsed from code; for example, consider HTML, Word, or Excel. It makes me wonder how screen readers work. Thinking out loud, it seems that the screen readers should let the applications (e.g., Word) handle their own parsing and presentation and obtain the data after that. Otherwise, the screen reader would have to reinvent many wheels, interpreting the code for all applications including all their versions, features, quirks, and platform integration issues - such a daunting and difficult task that it seems unlikely. But where do screen readers hook into the content? After it's output from the application but before it's an image for the screen (which could require OCR)? Ironically, I suppose PostScript or PDF could provide common interfaces.
- swiley 10y agoIt's not that there's code, it's that word and HTML are document formats while PostScript/PDF is a vector graphics format. Generally, if you remove the formatting tags from HTML or word (leaving just the text) the characters are extremely likely to be in the same order as the rendered document (sans things like running head and page numbers). Furthermore, HTML and word both explicitly delimit things like paragraphs while PostScript just changes the drawing position.
- contravariant 10y agoSure all formats require some parsing to get to the content. However Tex, HTML, Word, Excel and languages like that were designed to allow people to format text and other data into a document. Therefore they generally make text and other information appear sequentially, and separate display logic from content. PostScript and PDF have no such separation, the are the display logic. They are fully fledged programs that list the position, size, font, colour, of every symbol on every page, if you're lucky in a vaguely logical order, but there's no reason it should be. If the files were written by a human you may have some hope of extracting some of the content, but almost nobody writes PDF or PostScript by hand any-more.
- CJefferson 10y agoI think the difference with word, excel and Tex is that because they are editable, the content must be stored in a way where the flow is expressed, and each part can be clearly broken into its parts. With PDFs, all you have is the position of lines and characters on the page. There is no flow, ordering, or semantics.
- deleted 10y ago[deleted]
- witty_username 10y agoWell, less can extract the text; so I don't see why that's an issue.
- vsl 10y agoIt’s only an issue with crap PDFs (it’s possible to omit text information or obfuscate it to the point you copy & paste garbage out of them; pdfTeX-created PDFs are of course fine).
- CJefferson 10y agoIt doesn't work often for multi column pdfs or tables, ligatures are usually misparsed, and maths is just destroyed.
- noamyoungerm 10y agoIt's totally possible (and a relatively frequent occurrence) to have pdfs where the order of characters in the code has no relationship at all to how those same characters are laid out visually on the page. Anything marginally more complex than a series of paragraphs with no formatting at all basically requires you to render out the whole pdf and figure out the order that you are actually supposed to read the characters in.
- tripzilch 10y agoYup. For instance, Word's PDF output, has an absolutely positioned textbox for every word (and sometimes sub-word). This is for kerning purposes. If you want your original text back, you're going to need some OCR-like preprocessing and heuristics to guess what textboxes belong to the same line. If you have multiple columns, good luck distinguishing them from accidental rivers. It's not impossible, but I wouldn't know immediately what tools get this most right. And it's always a lossy operation going back and forth.