3 ms·
A perfectly-marked-up HTML document only solves the problem for text expressed in HTML. It doesn’t solve any other format that people publish (even today we ha
by makecheck 7y ago
A perfectly-marked-up HTML document only solves the problem for text expressed in HTML. It doesn’t solve any other format that people publish (even today we have links to PDF and Word and Excel, etc.). And of course there is meaningful information in formats that are not documents at all, like images. Some of those images are pictures of text that we might want to read.
It requires some investment to come up with tools for analyzing files. I’m wondering if the tools would have been as sophisticated if documents had lowered their barriers. For example, how do you justify developing a machine-learning model to look for more, if it seems all your documents are already semantically tagged with the details that were important to somebody?