4 ms·
Wouldn't web scraping be possible by taking screenshots of the rendered pages and then reading them with OCR?
by svdr 3y ago
Wouldn't web scraping be possible by taking screenshots of the rendered pages and then reading them with OCR?
- ekianjo 3y agoprobably very inefficient as it would depend on layout a lot too
- cush 3y agoAs inefficient as parsing heap snapshots?
- brigadier132 3y agoMuch more
- deleted 3y ago[deleted]
- spaniard89277 3y agoYou'll be spending resources on LLMs like crazy. Possible but very messy IMO.
- anamexis 3y agoYou don't need LLMs for OCR.
- ekianjo 3y agoOCR does not get you the names of the classes in a DOM
- spaniard89277 3y agoNo but maybe you want to do something with the ocr output.
- simonw 3y agoIf you just want the text there are other ways to do that. You could dump out document.body.innerText for example - here's how to do that with https://shot-scraper.datasette.io/en/stable/javascript.html https://shot-scraper.datasette.io/en/stable/javascript.html shot-scraper javascript youtube.com 'document.body.innerText' -r Output: https://gist.github.com/simonw/f497c90ca717006d0ee286ab086fbdde https://gist.github.com/simonw/f497c90ca717006d0ee286ab086fb... Or access the accessibility tree of the page using https://shot-scraper.datasette.io/en/stable/accessibility.html https://shot-scraper.datasette.io/en/stable/accessibility.ht... shot-scraper accessibility youtube.com Output here: https://gist.github.com/simonw/5174380dcd8c979af02e3dd74051a9c9 https://gist.github.com/simonw/5174380dcd8c979af02e3dd74051a...
- lelandfe 3y agoOf course, if the document is using the outline in unexpected ways, you'll run into trouble. Consider Facebook infamously splitting "Advertisement" into multiple spans to avoid tripping ad blockers.
- michaelt 3y agoAlthough you'd imagine screenshots would be easy to OCR reliably, it's not guaranteed to get everything correct. It's not like you can rely on a dictionary to confirm you've correctly OCRed a post by "@4EyedJediO" - who knows if that's an O or a 0 at the end? And if you're OCRing the title and view count of a youtube video, for example, you've got to take the page layout into account because there's a recommendations sidebar full of other titles with different view counts.
- plorntus 3y agoI guess you'd get better results if you knew the font the site uses (which in many cases you could figure it out pretty quickly) or even just override every font with your own.
- berkle4455 3y agoMuch of the content worth scraping isn't rendered on the screen.
- zffr 3y agoDo you have any examples? I haven’t experienced this myself
- throwawayadvsec 3y agoURL, images, stuff shown after you click on a button...
- is_true 3y agoYes, it's possible. We do this for TV shows.