4 ms·
Great article on the WARC file format. I really appreciate the efforts of the ArchiveIt team in developing and maintaining such a crucial tool for web archiving
by kburman 3y ago
Great article on the WARC file format. I really appreciate the efforts of the ArchiveIt team in developing and maintaining such a crucial tool for web archiving.
I'm trying to understand why the WARC format was made the way it is. Can someone explain the main reasons behind its design? Specifically, I'm curious about:
- Why not use a regular folder and file system with separate metadata files? It seems simpler and more user-friendly.
- Why not use SQLite?
- SnowflakeOnIce 3y agoAs to the your second question: WARC is based on an older format that was started in the 1990s, before SQLite existed.
- kburman 3y agoThanks for that info! Didn't realize WARC predated SQLite. Makes me wonder, are there any modern updates that could enhance WARC, or is it only format for archiving?
- makeworld 3y agoThe main update to WARC now is WACZ. https://specs.webrecorder.net/wacz/1.1.1/ https://specs.webrecorder.net/wacz/1.1.1/
- TheTechRobo 3y agoYeah, the format really needs an update. For starters, WARC only officially supports HTTP/1.1. Webrecorder has started faking HTTP/1.1 data in WARC files in order to save other versions, but I don't think faking data is great for an archival format, especially if it isn't standardized.
- egh 3y agoIt's a pretty simple design, and it's based on the ARC format (https://archive.org/web/researcher/ArcFileFormat.php https://archive.org/web/researcher/ArcFileFormat.php) which is even simpler. In response to your questions, here's my take (as somebody who used to work on web archiving). 1. Two reasons: First, many files are harder to manage. WARC files might contain hundreds or thousands of files. It's easier to manage big groups of files that are roughly the same size. Both for humans, and, at least in the past, for the file systems themselves. Second, once you break them up into files, what do you name the files? If you give them a name unrelated to the URL that was fetched, what is the advantage? If you name them based on the URL, suddenly you have a problem of mapping a URL to a legal file name, which can vary based on the file system. This would be a huge headache. 2. Yes, it predates SQLite, but also, why would you use sqlite? That's adding a huge amount of complexity. Is SQLite even good at storing big binary blobs? Additionally, because of the clever way that WARC files are gzipped, each piece of the WARC file is gzipped individually, which allows random access into the file for reading enclosed content in a compressed file without needing to read the entire WARC file.
- Retr0id 3y ago> Is SQLite even good at storing big binary blobs Nope! SQLite is good for lots of small-ish blobs (kilobytes), but once you start getting into the megabyte range, less so. There's also currently a hard upper limit blob size of 2GiB.
- kburman 3y agoThanks for the insights, egh! It's clear now why SQLite wouldn't be ideal for this purpose. Also, the point about URLs not always being valid filenames really makes sense.
- abracadaniel 3y agoTo add to this, WARC.gz files are also concatenated gzip records, so you can read any record by starting a decompression at a known offset. This gives you the access time of a file with the efficiency of having many many records only taking up one file.
- nikisweeting 3y agoWACZ also extends this functionality to allow streaming archives off a server without having to request the whole file to get one page. https://replayweb.page/docs/wacz-format https://replayweb.page/docs/wacz-format
- marginalia_nu 3y ago> - Why not use a regular folder and file system with separate metadata files? It seems simpler and more user-friendly. Filesystems don't generally deal well with having billions of files. Like you can work around some issues with deep directory structures, but even then it's not very effective and many cases you'll run out of inodes before you even reach a single billion. This is not a use-case filesystems in general tend to optimize for. > - Why not use SQLite? The main reason WARC looks like it does is to be as recoverable as possible if the crawler hard-crashes. There are much fewer things that can go wrong compared other formats. Records are only stored in exactly one location. Data corruption, write faults and other error states are all easy to detect and reason about. You have to work really hard to mess up a WARC file beyond being able to salvage most of it.
- kburman 3y ago> The main reason WARC looks like it does is to be as recoverable as possible if the crawler hard-crashes. I disagree on the recovery part. SQLite being ACID compliant, arguably offers better recovery than the WARC format.
- marginalia_nu 3y agoWARC is also resistant to hardware errors though, and has this property without requiring constant fsyncing. A CSV file is arguably harder to recover by comparison.
- kburman 3y ago> WARC is also resistant to hardware errors though, and has this property without requiring constant fsyncing. I just downloaded a sample WARC file to check for any checksums to detect bit rot, but I couldn't find one. Can you share any resources I can explore to understand how WARC is more resistant to hardware errors compared to other file formats?
- marginalia_nu 3y ago
- michaelt 3y ago> Why not use a regular folder and file system with separate metadata files? It seems simpler and more user-friendly. First of all, no matter how easy it is to extract the files, the results will be kinda user unfriendly - unless someone goes through the archived html files updating all the URLs for images and javascript and so on, it'll probably end up looking pretty broken. For users who just want to view a few files, Wayback Machine is a better tool for the job. With that said, some situations where warc is helpful include: * If a website had a page at www.example.com/foo and a file at www.example.com/foo/bar.txt you can't express that in a regular filesystem, as you can't have a file and a directory with the same name. * If a website uses some absurd 2000 character URL like https://s3.eu-west-2.amazonaws.com/document-api-images-live.ch.gov.uk/docs/D-WKi1U-gH-GK5_ZAteBCDH8U7KRu2oGH5riOkQkfy0M/application-pdf?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIAWRGBDBQ3JAGVZSXW%2F20240124%2Feu-west-2%2Fs3%2Faws4_request&X-Amz-Date=20240124T211617Z&X-Amz-Expires=60&X-Amz-Security-Token=IQoJb3JpZ2luX2VjEML%2F%2F%2F%2F%2F%2F%2F%2F%2F%2FwEaCWV1LXdlc3QtMiJHMEUCIAH8WTsMW%2Fr%2Bd4bDE1z%2FcEOtTnZ45BG%2BU2Dobz3%2FtpcWAiEAgN4c37JFBVlCvRUj5cRq1Zycuevq5w5PwzB0qu%2FcVsoquwUIexAEGgw0NDkyMjkwMzI4MjIiDFL6RhazU9lNl3hJpyqYBXstsc8y4MsNNSBiwkNcxeLNJXip8N9DyUs3X37F%2Bo0x%2BTKKm0f2kJ2%2FXxAijyiQVzNrL9luVcC3T7jeoPneZHm1qnEuVSitp300jCU7KRu2oMq3%2BDbGR1QUMp1QypNXJVCmyax5uCIzkB30XSwoQ2M57J3jSb4l9GC18UVAqVGsWralDgeuHNMneeiRQhJrI6zkVYw28kRhgeDl9tG%2F%2B29Ky3fr%2BEBnZSmKAjRL0R5wS37sRNVM3hofweHqEtKE69ai469OH0yIL13K4UwMaVUgpfLo24NPR9liLfiy7fTMAoC8SJhAfxwRc9c%2BnRfmY5RXS35c0dwaizYQJ6qaWGFq3ZUOgUZRIuxEOTSC%2BNoTHwduW9ZURbEd9AIEKrfMzM8SVo3JX6I3ZbpIlmbz87OTnWUQVb8BTLA%2F%2BEKAhJ9u%2BCHQbJ2Z6xfqssPAyTid859ExhRrUhsV4KB0QeaQJ5vKVLAT2vO2Oyd%2FATab5lMQLILbJB8NK2x64x8jQaYljswQnX6YR8IZr%2BI%2BwWo8HF%2FsQkgfI%2BHfJV4Wn%2FsGMB5jrfXm519lbogWVRyVx65%2FqKv8DbPvIYoz1tztfmyhk5PAiA9ugi4mm%2Fqxj2QSM08wcyU9wKMKJ0N8F1ABgK8H0FV%2F9mr9y5Mf5I5RCLQ%2B2qqjjQaiypC%2Fzf%2FAZEvGZlrrqmcnvCdDSJpiEMhr4MUQmTR5KmBBV5knnyIC9%2BLvVCQWqzgSxU5U%2BX4dStCK9nQwQX76Vrxu9Mgcl%2FQT1r3ii2LdfTjPRjmZ%2BY%2BTyTr%2FDVwYl%2Bmf4ky8wnJMpqMiO4Br5WEJIiuPjxWFdA4%2BuKKaKJot5XAdHLNvu8mwZ5G8NwJj6Keebomq3r1awjxPxglUqj9gwoejDIYwt6XFrQY6sQG7ggh2Xrja38H%2ByRtpBN6ZCbxCuSY6%2BN0ZGFHHg9r%2FPqcdP0Uguy%2B06X78ULxk6dEitIckzaXlLyfkexOkmDToxP5gsU7g%2FUJvZAx9oqfU51zwvteeQwLqmjvB4EG4wZozziRMaB9L0pPTi8w5ZqanBU9SxHZNrrRHNDphioNGXHo9foXlrDHK3iRXkgftY2P0m78ujp0J2ZZbDMbrxoY1OjLly54dRVX0tu2mVoWWW9s%3D&X-Amz-SignedHeaders=host&response-content-disposition=inline%3Bfilename%3D%22companies_house_document.pdf%22&X-Amz-Signature=9cb107709fbc76576e5d2d4c9f14a3b61982dee78831dd45af03317da9040414 https://s3.eu-west-2.amazonaws.com/document-api-images-live.... you don't end up with a filesystem-breaking filename. * If www.example.com embeds an image from exampleusercontent.com you can capture the image in the same archive file. * If for some reason you want to store daily copies of www.cnn.com in the same archive for comparison purposes - you can. * It lets you store headers, 300 redirect messages, case-sensitive filenames, and all that sort of stuff. * And it's an extremely simple format - basically human readable. So if you think you're archiving for the super-long-term and want to make really conservative choices, you can be pretty confident in plain text.
- fforflo 3y ago> - Why not use SQLite? https://github.com/Florents-Tselai/WarcDB https://github.com/Florents-Tselai/WarcDB