5 ms·
At this point, the easier route to getting a text-based browser to support something like this would be creating a new one based on WebKit/Blink. It would prob
by halter73 7y ago
At this point, the easier route to getting a text-based browser to support something like this would be creating a new one based on WebKit/Blink.
It would probably need to get the engine to draw to a fake screen buffer and run an OCR algorithm over that. And even if the OCR and layout worked well, there would be a lot other work necessary to get reasonable text based interactivity, though that's probably partially solved by projects like Vimium. Some interactions like dragging would likely never be supported.
Not easy at all, but somehow more reasonable than updating lynx to support all of today's new technologies. I wonder if anyone's already tried something like it.
- halter73 7y agoI just found Browsh[1] which seems to be basically this but with Gecko. It seems they don't need to do OCR which makes sense thinking about it more. Cool stuff. Haven't gotten to try it yet. [1] https://www.brow.sh/ https://www.brow.sh/
- tombh 7y agoI did! See brow.sh It uses Firefox as a backend so can in fact run Wasm. However it doesn't use OCR, it uses the DOM to get precise coordinates for text nodes, this recreating a pure text representation of the page, using nothing but spaces and carriage returns for "formatting".
- colordrops 7y agoThe vim browser demo uses pixels drawn to canvas though.
- aasasd 7y ago> OCR Hey, dial back the blasphemy there. HTML is not PDF, it's usually made from text in the first place. I'd guess you could force all text to use a monospace font, with fixed measures and line-height, and limit the width of the page. Then mostly dump the resulting text arrangement into the terminal. Now, layouts from the various elements and CSS are probably a lot trickier, but snapping all margins and padding to multiples of a symbol's size should go a long way. It seems that this could even be embedded at different levels in the browser: the layout engine or just the user's JS. (If JS can obtain the exact layout of text lines and elements―likely not, though, especially in forms. Maybe via devtools.)
- _pmf_ 7y ago> HTML is not PDF, it's usually made from text in the first place. I don't know, man. Most websites would be smaller if they were replaced by a HD video of someone reading the contents.
- jsjohnst 7y ago> Hey, dial back the blasphemy there. HTML is not PDF, it's usually made from text in the first place. Hey, dial back the blasphemy there. PDF is not an image, it’s usually[0] made from text in the first place. As someone who has a bunch of experience both creating and parsing PDFs, it’s definitely very doable to extract the text content and render it in a similar way on a terminal. Yes, parsing the PDF format is much more painful than average HTML, but these days there’s libraries commonly available to assist. [0] unless the PDF is a famous redacted DOJ document, then it’s a poorly scanned collection of image crammed into a PDF container.