3 ms·
I think they actually have a point. Parsing HTML isn’t trivial: aside from bad/invalid HTML (think: missing opening/closing tags, quotes, etc), there’s also a l
by koprulusector 4y ago
I think they actually have a point. Parsing HTML isn’t trivial: aside from bad/invalid HTML (think: missing opening/closing tags, quotes, etc), there’s also a lot of content that requires javascript to render in the first place, for example, which means the page needs to be rendered and have access to window and DOM, etc.
Then, of course, you have embedded objects, such as iframes, or that aren’t text that traditional parsing can’t easily identify/extract, even if everything is rendered OK. For example, a video, animations, interactivity, etc. Another contrived example which illustrates the point would be a site that uses images for buttons/links rather than text, or even non-semantic HTML such as a <div> and click handlers as "links", or buttons with <a> tags.
I can think of several other examples.
A "screenshot API" enables automation to capture "picture of a group of people celebrating" distinct from advertisements that might also appear on the page without the need to dig through CSS selectors/classes, domain name/query parameters, and handles cases where the image might be base64 embedded directly. Another example might be to simply/easily extract data from a Google Sheets/Excel table, which itself might include embedded images or non-text/HTML objects. Such an API could help accessibility by enabling screen readers for sites that weren't built to be accessible.
I learned from this post that you can tap/copy objects from photos in iOS 16! I just did this for the first time and it’s CRAZY, using a photo of my basement and a bunch of shoe boxes, tools, etc. I pressure tapped the group of shoe boxes and hit copy, pasted into the Bear app (a markdown editor for Apple devices), and it pasted just the shoe boxes, perfectly clipped around the edges as if I’d used photoshop!
I think ultimately, the point is that such an API which uses object detection, image-to-text, sentiment analysis, etc., on the backend could make trivial the tasks and edge cases that today require non-trivial effort and time, and could enrich the data prior to its retrieval.
- pwdisswordfish9 4y ago> Parsing HTML isn’t trivial: aside from bad/invalid HTML (think: missing opening/closing tags, quotes, etc), there’s also a lot of content that requires javascript to render in the first place, for example, which means the page needs to be rendered and have access to window and DOM, etc. Double standard. If you're going to make a fair comparison, then you need to compare like with like; you need to compare the subset of things about e.g. HTML that give you what you can also get with a screenshot. It makes no sense to hold the performance penalty of script execution against browser runtimes when (a) you don't have to execute any scripts to effect anything that gives you parity with a static image, and (b) you can't with static images do anything like what executable scripts enable. And whether or not parsing HTML is trivial (which is debatable), it's still not strictly greater than the computational resources that are needed for the kind of computer vision and widgetry that lets you e.g. select the text in a screenshot...