4 ms·
How do people archive sites 1:1 using tools like Playwright? I’ve tried to screenshot things (looks weird) and pull the page content (have problems viewing arti
by server_man3000 3y ago
How do people archive sites 1:1 using tools like Playwright? I’ve tried to screenshot things (looks weird) and pull the page content (have problems viewing articles like medium).
- bsnnkv 3y agoIt's ultimately a cat and mouse game with many websites actively trying to sabotage archival efforts. Unless this is a core functionality in something you're working on, most people will be better off using the SavePageNow API from archive.org and integrating with that. This is what I ultimately ended up doing for one of my projects.[1] [1]: https://lgug2z.com/articles/notado-07-2023-update/ https://lgug2z.com/articles/notado-07-2023-update/
- ajvs 3y agoWebsite owners can request their site be blacklisted for archival, so this doesn't work for all websites.
- linusg789 3y agoArchive.today is another popular site that doesn't respond to blacklist requests.
- bsnnkv 3y agoYeah this is true, but it works for enough websites to be a meaningful option, and it is almost always going to work better than something you have home-rolled (unless your core product is a direct competitor or something, in which case all bets are off ;)) A good is example of this is Pinboard which claims to offer website archiving. A friend has over 100,000 links saved (with an archival account) there. When we spent a few minutes looking at the archive links for those items a few weeks ago, we couldn't find a single correct, working, accessible archive from the most recently archived links (listed as archived 5 weeks ago, so also not up to date).
- e12e 3y agoI wonder if Firefox "reader mode as a utility" might be a viable alternative for Pinboard like "content oriented" archiving? https://github.com/mozilla/readability https://github.com/mozilla/readability
- jacobwilliamroy 3y agoMost of the time a client just cares about the information, not the typesetting or the layout, and in that case most of the data can just be pulled from the web inspector (downloading files). I've yet to have a client ask me to also copy the typesetting and layout so I never learned how to do that.
- tw4l 3y ago(Disclaimer: I work at Webrecorder) Our automated crawler browsertrix-crawler (https://github.com/webrecorder/browsertrix-crawler https://github.com/webrecorder/browsertrix-crawler) uses Puppeteer to run browsers that we archive in by loading pages, running behaviors such as auto-scroll, and then recording the request/response traffic in the WARC format (by default in Webrecorder tools, then packaged into a portable WACZ file: https://specs.webrecorder.net/wacz/1.1.1/ https://specs.webrecorder.net/wacz/1.1.1/). We have custom behaviors for some social media and video sites to make sure that content is appropriately captured. It is a bit of a cat-and-mouse game as we have to continue to update these behaviors as sites change, but for the most part it works pretty well. The crawler also has some job queuing functionality, supports multiple workers/browsers, and is highly configurable to set timeouts, page limits, etc. The trickier part is in replaying the archived websites, as a certain amount of re-writing has to happen in order to make sure the HTML and JS are working with archived assets rather than the live web. One implementation of this is replayweb.page (https://github.com/webrecorder/replayweb.page https://github.com/webrecorder/replayweb.page), which does all of the rewriting client-side in the browser. This sets you interact with archived websites in WARC or WACZ format as if interacting with the original site. replayweb.page can run locally in your browser without needing to send any data to a server or can be hosted, including in an embedded mode. (edit: fixed typos)
- tw4l 3y agoWe experimented moving to Playwright but Playwright doesn't handle long-running browser sessions well, as the devs have (maybe rightly) prioritized its use for testing over archival use cases and want you to spin up a browser each time. For archival purposes, that doesn't work as well because we're not able to save the browser profile to retain cookies such as login credentials, so we've moved back to using Puppeteer for now.
- niam 3y agoAnswered my next question before I even asked it. As a next-next question: May I ask what y'all experienced with long-running browser sessions? Or what in particular led you to believe it was unfit for this purpose?