4 ms·
The Wayback backend is much simpler than one might imagine. The index lookup ("where is the data for this URL at this timestamp?") is called the CDX API and is
by bnewbold 9y ago
The Wayback backend is much simpler than one might imagine. The index lookup ("where is the data for this URL at this timestamp?") is called the CDX API and is pretty fast, given that it's basically looking up a line in a sorted many-terabyte text file. Slower are the processes that extract the original HTTP response from what is effectively a giant multi-GB tarball sitting on an archival-grade (not IOPS-optimized) spinning disk, and the code that parses HTML and/or javascript and re-writes "embed" URLs for playback. Another speed limit is that we serve everything from our datacenters in the bay area with no CDN for most content, which you'll really notice when connecting from outside North America.
The priority is to have as much data as accessible as possible, and for the same cost we can crawl and store far more web content on slow spinning disks than with SSDs or large RAM caches. Most RAM on storage nodes, which would otherwise be used as disk cache, gets used for derive tasks (like OCR or video conversion) or crawling/crunching tasks. That being said, if the service is so slow as to be unusable then there's no point operating it in the first place; hopefully we can get the latency a it lower.
- johansch 9y agoSo the primary reason for the multiple second response times is that many sequential megabytes have to be read from an HD in order to return an often quite tiny HTTP response data for a particular {url, timestamp}? That would explain the nearly constantly (slow) response time. (I was puzzled why it didn't seem to vary much depending on the time of day, etc.)
- bnewbold 9y agoWell, I oversimplified a bit. The "tarball" (WARC or ARC file) is a single large file with individually compressed HTTP response objects concatenated together. The index stores the byte offset and length (compressed) into this file, so an HTTP/1.1 range request can pull only the data needed. So in theory it's efficient. I don't work on that team directly and haven't looked in to exactly what the performance latency bottleneck is, sorry for the misleading response.
- johansch 9y agoOk. Well, I do hope you work on improving the response times. I think it would do wonders for the popularity of the wayback machine (and in the end, financial contributions).