7 ms·
Is there a reason why you'd want to save archival material in a proprietary format? Wouldn't it be better/easier to use wget with the `--warc-file` flag?
by eltondegeneres 2y ago
Is there a reason why you'd want to save archival material in a proprietary format? Wouldn't it be better/easier to use wget with the `--warc-file` flag?
- simongr3dal 2y agoIt is proprietary, but it's not like it's difficult to decode or interpret.
- deleted 2y ago[deleted]
- codetrotter 2y agoFTA: > Although Safari is only maintained by Apple, the Safari webarchive format can be read by non-Apple tools – it’s a binary property list that stores the raw bytes of the original files. I’m comfortable that I’ll be able to open these archives for a while, even if Safari unexpectedly goes away.
- deleted 2y ago[deleted]
- alexwlchan 2y ago1/ Why not wget? For this project I wanted a consistent file format for my entire collection. I have a bunch of stuff I want to save which is behind paywalls/logins/clickthroughs that are tricky for wget to reach. I know I can hand wget a cookies file, but that’s mildly fiddly. I save those pages as Safari webarchive files, and then they can drop in alongside the files I’ve collected programatically. Then I can deal with all my saved pages as a homogeneous set, rather than being split into two formats. Plus I couldn't find anybody who'd done this, and it was fun :D This is only for personal stuff where I know I'll be using Safari/macOS for the foreseeable future. I don't envisage using this for anything professional, or a shared archive -- you're right that a less proprietary format would be better in those contexts. I think I'm in a bit of a niche here. (I'm honestly surprised this is on the front page; I didn't think anybody else would be that interested.) 2/ Proprietary format: it is, but before I started I did some experiments to see what's actually inside. It's a binary plist and I can recover all the underlying HTML/CSS/JS files with Python, so I'm not totally hosed if Safari goes away. Notes on that here: https://alexwlchan.net/til/2024/whats-inside-safari-webarchive/ https://alexwlchan.net/til/2024/whats-inside-safari-webarchi...
- pvg 2y agoI didn't think anybody else would be that interested. 'Save the webpage as I see it in my browser' remains a surprisingly annoying and fiddly problem, especially programmatically, so the niche is probably a little roomier than you might initially suspect.
- diggan 2y ago> 'Save the webpage as I see it in my browser' remains a surprisingly annoying and fiddly problem Is it really? I remember hacking around with with JavaScript's XMLSerializer (I think) like 5 years ago and solved that for ~90% of the websites I tried to archive. It'd save the DOM as-is when executed. Internet Archive/ArchiveTeam also worked on that particular problem for a very long time, and are mostly successful as far as I can tell.
- pvg 2y ago90% feels like an overestimate to me but it's already quite poor, you wouldn't accept that for saving most other things. Another problem is highlighted in the piece - it's a hassle to ensure external tools handle session state and credentials. Dynamic content is poorly handled, the default behaviours are miserable (a browser will run random Javascript from the network but not Javascript you've saved, etc). There's a lot of interest in 'digital preservation' and perhaps one sign of how it's very much early days of the field - it's tricky to 'just save' the results of one of the most basic current computer interactions - looking at a web page.
- diggan 2y agoBut if you serialize the DOM as-is, you literally get what you see on the page when you archive it. Nothing about it is dynamic, and there is no sessions nor credentials to handle. Granted, it's a static copy of a specific single page. If you need more than that, then WARC is probably the best. For my measly needs of just preserving exactly what I see, serializing the DOM and saving the result seems to do just fine.
- buildbot 2y agoHow can you actually open and view a warc file? I've never found a good application for this, have I missed something obvious?
- diggan 2y agoLots of tools available, best index I've found of the ecosystem is this: https://wiki.archiveteam.org/index.php/The_WARC_Ecosystem https://wiki.archiveteam.org/index.php/The_WARC_Ecosystem Ultimately, this is the best viewer I've found so far: https://replayweb.page/ https://replayweb.page/
- tivert 2y ago> Lots of tools available, best index I've found of the ecosystem is this: https://wiki.archiveteam.org/index.php/The_WARC_Ecosystem https://wiki.archiveteam.org/index.php/The_WARC_Ecosystem It looks like there's a lot of tools for creating them, but not a a lot for viewing. What they really need is browser support, or at least an extension so a browser can open the files directly.
- jamesgeck0 2y agoThere is a browser extension. It can record WARC files, but also has a viewing interface identical to ReplayWeb.page. https://archiveweb.page/guide https://archiveweb.page/guide
- zie 2y agoSadly it's Chrome only it seems.
- cxr 2y ago> What they really need is browser support, or at least an extension so a browser can open the files directly That's probably the wrong thing. What browsers really need is a thin but standardized API that lets any third-party app that the user has installed on their machine to supply the content for various fetch/reads. You'd open the WARC in Firefox or Safari or whatever, but Safari et al wouldn't have any special understanding of the format. It would know that your app does WARCs, though, and then knock on the door and say, "Please tell me the content I should be showing here; I'll defer to you for any further "requests" associated with the file/page loaded in this tab—just tell me the content I should use for those, too."