3 ms·
Almost all formats would similarly require content to be parsed from code; for example, consider HTML, Word, or Excel. It makes me wonder how screen readers wo
by hackuser 10y ago
Almost all formats would similarly require content to be parsed from code; for example, consider HTML, Word, or Excel.
It makes me wonder how screen readers work. Thinking out loud, it seems that the screen readers should let the applications (e.g., Word) handle their own parsing and presentation and obtain the data after that. Otherwise, the screen reader would have to reinvent many wheels, interpreting the code for all applications including all their versions, features, quirks, and platform integration issues - such a daunting and difficult task that it seems unlikely. But where do screen readers hook into the content? After it's output from the application but before it's an image for the screen (which could require OCR)? Ironically, I suppose PostScript or PDF could provide common interfaces.
- swiley 10y agoIt's not that there's code, it's that word and HTML are document formats while PostScript/PDF is a vector graphics format. Generally, if you remove the formatting tags from HTML or word (leaving just the text) the characters are extremely likely to be in the same order as the rendered document (sans things like running head and page numbers). Furthermore, HTML and word both explicitly delimit things like paragraphs while PostScript just changes the drawing position.
- contravariant 10y agoSure all formats require some parsing to get to the content. However Tex, HTML, Word, Excel and languages like that were designed to allow people to format text and other data into a document. Therefore they generally make text and other information appear sequentially, and separate display logic from content. PostScript and PDF have no such separation, the are the display logic. They are fully fledged programs that list the position, size, font, colour, of every symbol on every page, if you're lucky in a vaguely logical order, but there's no reason it should be. If the files were written by a human you may have some hope of extracting some of the content, but almost nobody writes PDF or PostScript by hand any-more.
- CJefferson 10y agoI think the difference with word, excel and Tex is that because they are editable, the content must be stored in a way where the flow is expressed, and each part can be clearly broken into its parts. With PDFs, all you have is the position of lines and characters on the page. There is no flow, ordering, or semantics.
- deleted 10y ago[deleted]