4 ms·
I built this as a small utility within a larger project I'm working on these days. (contact me if you're curious or want to support). The WARC format is extrem
by fforflo 4y ago
I built this as a small utility within a larger project I'm working on these days. (contact me if you're curious or want to support).
The WARC format is extremely simple and yet so powerful. Most importantly though, there are already pebibytes of already crawled archives.
This is a fairly straightforward mapping of a .warc file to a .sqlite database. The goal is to make such archives SQL-able even in smaller pieces.
The schema I've come up with it's tailored around my requirements, but comment if you can spot any obvious pitfalls.
PS: I do believe that at some point .sqlite will become the defacto standard for such initiatives. Sure, it's not text... but it's pretty close.
- marginalia_nu 4y ago> PS: I do believe that at some point .sqlite will become the defacto standard for such initiatives. Sure, it's not text... but it's pretty close. What is the advantage of moving around .sqlite-files, over just loading the (compressed) WARCs into sqlite databases when you need them? I've been messing around with different formats for my own search engine crawls, and ended up with the conclusion that WARC is a pretty amazing intermediary format that weighs both the needs of the producer and consumer very well. I don't use WARCs now, instead something similar, but I probably will migrate toward that format eventually. WARC's real selling pint is that it's such an extremely portable format.
- uniqueuid 4y agoAnecdotal evidence, but I produced a medium-size crawl in the past (~20TB compressed). I used distributed resources with off-the-shelf libraries (i.e. warcprox etc.) and managed to get corrupted data in some cases where neither the length-delimited (i.e. offset + payload length) nor the newline-delimited (triple newlines between records) logics were valid any longer. Took me some time to build a repair tool for that. Sqlite has an amazing set of well-understood and documented guarantees on top of performance, there's a host of potential validation tools to choose from and you can even use transactions etc. So that alone seems like a great idea. What's more, you can potentially skip CDX files if you have sqlite databases (or build your own meta sqlite database for the others quickly).
- rengler33 4y agoIs there a forum or somewhere web crawlers hang out online? I'd love the learn about more sophisticated projects like this.
- uniqueuid 4y agoIn github issues of said projects, and at scientific web archival conferences. Although I'd absolutely welcome some sort of channel!
- smcnally 4y agoTopics include Issues and activities across projects. These topics are quite active, e.g. https://github.com/topics/crawling https://github.com/topics/crawling https://github.com/topics/web-scraping https://github.com/topics/web-scraping https://github.com/topics/web-archiving https://github.com/topics/web-archiving
- ikreymer 4y agoA bit late to this thread, but I think WARC is a reasonable format for raw HTTP traffic. We should definitely have better tools to ensure WARC files produced are valid, and that's one of the things we build at Webrecorder. Unless you're crawling really text heavy content, most of the WARC data is binary content that doesn't really need to be in a db. However, sqlite or database as replacement for CDX is an appealing option, where WARC files can remain static data at rest and derived data (offests, full-text search, can be put into a db. We are experimenting with a new format, WACZ, which bundles WARC files into a ZIP, while adding CDXJ and exploring sqlite as an option for full-text search. I agree that it's better to build on solid, existing formats that can be validated, especially when large amounts of data are concerned!
- fforflo 4y ago> What is the advantage of moving around .sqlite-files, over just loading the (compressed) WARCs into sqlite databases when you need them? The .warc spec is ideal. I'm not saying we replace it (ref. xkcd: standards). On top of what uniqueid said, "loading" is much slower and more cumbersome than it sounds. I'm not saying SQLite will replace text (maybe my aphorism sounded too firm). I'm saying that maybe along with the .warc.gz archives at rest, one could have .sql.gz files at rest as well. In other words: why not move ACID-compliant archives moving around?
- marginalia_nu 4y agoSeems like the benefit of sqlite is the sort of usecases where you maybe don't want to load everything you've crawled into a search engine, but want to be able to cherrypick the data and retrieve specific documents for further processing. Which is certainly a use case that exists, and indeed not really what WARC is designed for.
- traverseda 4y agoGreat for most end user facing applications though.
- marginalia_nu 4y agoEnd-user facing applications usually don't consume website crawls, do they? That's impractical for many reasons, the sheer size alone being perhaps the biggest obstacle. If you want to do something like have an offline copy of a website, ZIM[1] is a far more suitable format as it's extremely space-efficient and also fast. [1] https://docs.fileformat.com/compression/zim/ https://docs.fileformat.com/compression/zim/
- mynameismon 4y ago> What is the advantage of moving around .sqlite-files, over just loading the (compressed) WARCs into sqlite databases when you need them? I suppose there could be made a case for easy extension: You don't need to change the entire spec to add another table in the SQLite database, maybe containing other metadata. > WARC's real selling pint is that it's such an extremely portable format. I mean, so is SQLite: it is also apporoved as a LoC archival method. (See SQL Archive [0]) [0]: https://www.sqlite.org/sqlar.html https://www.sqlite.org/sqlar.html
- nlohmann 4y agoHave you every played with SQLite virtual tables (https://sqlite.org/vtab.html https://sqlite.org/vtab.html) - they could allow to provide an SQLite interface while keeping the same structure on disk. Though it requires a bit of work (implementing the interface can be tedious), it can avoid the conversion in the first place.
- fforflo 4y agoGood point. Actually CommonCrawl provides Parquet files for their archives too. And there's this vtable for Parquet extension. https://github.com/cldellow/sqlite-parquet-vtable https://github.com/cldellow/sqlite-parquet-vtable But for my use case virtual would be too complicated.
- m_ke 4y agoDuckDB would probably be a way better option and works amazingly well on top of parquet (https://duckdb.org/docs/data/parquet https://duckdb.org/docs/data/parquet)
- fforflo 4y agoThen again, do you need virtual tables? The .warc structure won't change, so the tables won't change. But you can have SQL views defined instead for common queries.