6 ms·
Show HN: WarcDB: Web crawl data as SQLite databases
- uniqueuid 4y agoSince this is pretty new, some background: WARC is a file format written and read by a rather small but specialized set of web crawler tools, most notably the internet archive's tooling. For example, its java-based crawler heretrix produces warc files. There are a couple of other very cool tools, such as warcprox, which can create web archives from web browser activity by acting as a proxy, pywb which play back the same to make archived versions browsable, and related libraries [1](shoutout to Noah Levitt, Ilya Kreymer and collaborators for building all this). The file format itself is an iso standard, and it's very simple: Capture all http headers sent and received, and all http bodies sent and received, and do simple de-duplication based on digests of the body. There is a companion format, CDX, which builds indexes from warc files (which in turn are just concatenated records, so rather robust). Although all of this is great, I worry a bit about where we're heading with QUIC / udp-based protocols, websockets and other very involved protocols which ultimately make archival much harder. If there's anything you can do to help these (or other) tools to keep our web's archival records alive and flowing, please do so. [1] https://github.com/internetarchive/warcprox https://github.com/internetarchive/warcprox
- fforflo 4y agoThe web archiving community is surprisingly small and fragmented (in terms of tools) given its impact. Thankfully the .warc format looks pretty powerful and standard for the web we have so far (which is a lot! ). Now with the new protocols, dunno maybe its too soon to worry? Then again, maybe its an IPv4 / IPv6 analogy. > There is a companion format, CDX, which builds indexes from warc files (which in turn are just concatenated records, so rather robust). Good point. I's planning of combining this fact with the ATTACH option that SQLite has - allowing to query multiple database files [0] [0] https://www.sqlite.org/lang_attach.html https://www.sqlite.org/lang_attach.html
- uniqueuid 4y agoOh hi, thanks for building this! I haven't had the chance to play with it, but my hunch is that sqlite for warc can fill a great niche and would be much more portable (and probably performant). Allowing multiple DB files is a great idea, since that fundamentally enables large archives, cold-storage and so on.
- pstuart 4y agoSwarm and Union should be of interest to you: https://www.sqlite.org/swarmvtab.html https://www.sqlite.org/unionvtab.html
- ma2rten 4y agoCommon Crawl is also in WARC format.
- lijogdfljk 4y agoWould a WARC format reduce effort needed to make Reader-like programs? Ie strip pages of HTML cruft, leaving you with text, images, etc - the content?
- uniqueuid 4y agoNot at all. A WARC file gives you exactly what you would see on the wire, or in the network inspector tab of your browser. It does nothing to the content, and that's the point. The only thing you gain (and that's very important for other reasons as well) is an immutable ground truth to work from when creating the reader view of a given article.
- lijogdfljk 4y agoGotcha - yea i was hoping maybe it snapshotted the HTML or some such, side stepping some issues long dynamic text or JS shenanigans
- simonw 4y agoI've found the Readability.js library to be really good for that - here's my recipe for running it as a CLI: https://til.simonwillison.net/shot-scraper/readability https://til.simonwillison.net/shot-scraper/readability
- deleted 4y ago[deleted]
- fforflo 4y agoI built this as a small utility within a larger project I'm working on these days. (contact me if you're curious or want to support). The WARC format is extremely simple and yet so powerful. Most importantly though, there are already pebibytes of already crawled archives. This is a fairly straightforward mapping of a .warc file to a .sqlite database. The goal is to make such archives SQL-able even in smaller pieces. The schema I've come up with it's tailored around my requirements, but comment if you can spot any obvious pitfalls. PS: I do believe that at some point .sqlite will become the defacto standard for such initiatives. Sure, it's not text... but it's pretty close.
- marginalia_nu 4y ago> PS: I do believe that at some point .sqlite will become the defacto standard for such initiatives. Sure, it's not text... but it's pretty close. What is the advantage of moving around .sqlite-files, over just loading the (compressed) WARCs into sqlite databases when you need them? I've been messing around with different formats for my own search engine crawls, and ended up with the conclusion that WARC is a pretty amazing intermediary format that weighs both the needs of the producer and consumer very well. I don't use WARCs now, instead something similar, but I probably will migrate toward that format eventually. WARC's real selling pint is that it's such an extremely portable format.
- uniqueuid 4y agoAnecdotal evidence, but I produced a medium-size crawl in the past (~20TB compressed). I used distributed resources with off-the-shelf libraries (i.e. warcprox etc.) and managed to get corrupted data in some cases where neither the length-delimited (i.e. offset + payload length) nor the newline-delimited (triple newlines between records) logics were valid any longer. Took me some time to build a repair tool for that. Sqlite has an amazing set of well-understood and documented guarantees on top of performance, there's a host of potential validation tools to choose from and you can even use transactions etc. So that alone seems like a great idea. What's more, you can potentially skip CDX files if you have sqlite databases (or build your own meta sqlite database for the others quickly).
- jbverschoor 4y agoI like that you're logging responses instead of just the result / payload. Reminds me of some idea I had of using something like queue as an intermediary between a webserver and the backend. I don't exactly remember my reasoning anymore right now
- tepitoperrito 4y agoIt'd be neat to extend the warc format and tooling to support cached http responses for things like REST endpoints. Then you could make sure everything you did in a session is recorded for later use. From the specification it would appear fairly straightforward once an approach was chosen... Here's the relevant extract from the warc spec that informed my difficulty estimate - "The WARC (Web ARChive) file format offers a convention for concatenating multiple resource records (data objects), each consisting of a set of simple text headers and an arbitrary data block into one long file." Edit: Upon 2 minutes of reflection I think the way to go for what I'm envisioning is some kind of browser session recording -> replayable archive solution.
- uniqueuid 4y agoAlthough this is a nice idea, it's extremely difficult to get it completely right. Consider a SPA where navigation happens via xhr or similar requests and updates are json that's patched into the DOM. Even browsers have a hard time figuring out how to make this a coherent session. Now with warc, you get a single record per transfer, i.e. every json file, every image, every css is an individual record. It's completely up to the client/downstream tech to re-assemble this into a coherent page. If you want to go down that road, my best suggestion would be to start with a browser's history - that's probably the most solid version of a session that we have right now.
- nikisweeting 4y agoArchiveWeb.page + ReplayWeb.page are able to achieve what you're describing.
- TedDoesntTalk 4y agoDid you add this to https://github.com/dhamaniasad/WARCTools https://github.com/dhamaniasad/WARCTools ?