4 ms·
A point of great pain for me in the use of HTTP range headers is that they are absolutely incompatible with HTTP compression. This boils down to an ambiguity f
by apignotti 4y ago
A point of great pain for me in the use of HTTP range headers is that they are absolutely incompatible with HTTP compression.
This boils down to an ambiguity found in the early days of HTTP: when you ask for a byte range with compression enabled, are you referring to a range in the uncompressed or compressed stream?
The only reasonable answer, in my opinion: a range of bytes in the uncompressed version, which is then compressed on-the-fly to save bandwidth.
I suspect the controversial point was that originally HTTP gzip encoding was used via pre-compression of the static files. On-the-fly encoding was probably too computationally expensive at the time.
Anyway, the result of this choice done decades ago is currently preventing effective use of the ranged HTTP requests. As an example, it is fairly likely that chunks from the SQLite database can be significantly compressed, but both HTTP servers and CDN refuse to compress "206 Partial Content" replies.
Interestingly, another use case of HTTP byte ranges is video streaming, which is not affected by this problem since video is _already_ compressed and there is no redundancy to remove anymore.
- matharmin 4y agoI think another issue with compressing partial content is that those compressed responses cannot be cached efficiently - it would have to be compressed on-the-fly for every range requested. And while compression is not as computationally expensive as it used to be, it does still add overhead that could be more than the bandwidth overhead of uncompressed data. There should be a workaround for this using a custom pre-compression scheme, instead of relying on HTTP for the compression. Blocks within the file will have to be compressed separately, and you'll need some kind of index mapping uncompressed block offsets to compressed block offsets. Unfortunately there doesn't seem to be a common of-the-shelf compression format that does this, and it means that you can't just use a standard SQLite file anymore. But it's definitely possible.
- apignotti 4y agoI see the caching argument only partially. It applies to any ranged request, non just compressed ones, and there are plenty of ways for origin servers to describe what and how to cache content. The origin may explicitly only allow caching when the access patterns are expected to likely repeat.
- matharmin 4y agoUpdate: This is an example of a compression format that allows random access to 64KB chunks, and is compatible with gzip:https://manpages.ubuntu.com/manpages/kinetic/en/man1/dictzip.1.html https://manpages.ubuntu.com/manpages/kinetic/en/man1/dictzip...
- punnerud 4y agoWhy do you need HTTP compression, when all of SQLite is compressed? Will it give any value? Would think HTTP2.0 is of greater value, to save TCP round trip time.
- apignotti 4y agoI do not think SQLite compresses data by default. Can you point to any documentation suggesting otherwise?
- nbevans 4y agoSQLite isn't compressed. There is a commercial offering from the creators of SQLite that adds compression but its rarely seen in the wild to be honest.
- rjmunro 4y agoSQLite isn't compressed by default, but there are extensions like https://phiresky.github.io/blog/2022/sqlite-zstd/ https://phiresky.github.io/blog/2022/sqlite-zstd/ that offer compression. It would be interesting to see how well that extension works with this project.
- blueflow 4y ago> This boils down to an ambiguity found in the early days of HTTP The RFC is clear on this, different Transfer-Encoding is only for the transfer and does not affect the identity of the resource, while different Content-Encoding does affect the identity of the resource. Its not sure who did it first, but either Browsers or Web Servers started using the Content-Encoding header (which means, "this file is always compressed" like a tar.gz, user agent not meant to uncompress) with the meaning of Transfer-Encoding (which means, the file is not compressed, compression is just applied for the sake of the transfer and the user agent needs to uncompress first). This is a violation of the HTTP spec. This fuckup resulted in big confusion and now requires an annoying amount of workarounds: - E-Tag re-generation [1] - Ranges and Compression are not usable together - Browsers's "Resume Download" feature removed - RFC-compliant behavior being a bug [2] - Some corporate proxies are still unpacking tarballs on the fly, breaking checksum verification Contrast this to eMail, which has the Encodings correctly sorted out, while using the same RFC822 encoding technique as HTTP. [1] https://bz.apache.org/bugzilla/show_bug.cgi?id=39727 https://bz.apache.org/bugzilla/show_bug.cgi?id=39727 [2] https://serverfault.com/questions/915171/apache-server-causes-tar-gz-files-to-download-as-uncompressed-tarballs-still-nam https://serverfault.com/questions/915171/apache-server-cause... tl;dr: HTTP Standard is not being implemented correctly