4 ms·
A long time ago, I wrote something similar ( https://github.com/htcat/htcat https://github.com/htcat/htcat) to assist with Heroku's efforts to speed up moving t
by fdr 5y ago
A long time ago, I wrote something similar (
https://github.com/htcat/htcat https://github.com/htcat/htcat) to assist with Heroku's efforts to speed up moving the tar formatted application releases around. It's pretty old, it doesn't integrate tar archival itself, it probably can stand improvement. Or given its small size, even a rewrite.
My favorite hack in there (that also made it work with pre-signed S3 urls) was not using the HEAD method as is customary to determine object sizes, but instead doing a regular "GET" that, for small files, would execute on its own...but for larger would simply be abruptly closed by htcat once it reached the bytes that had since been fetched in parallel by a range-based request sent immediately afterwards. The goal was to have htcat not offer a penalty on small files so it could be used on blended workloads without thinking.
It also found a bug in S3's range implementation. We had a problem with some object or other, I wrote in about it, and was told that upon investigation a bug had been fixed. No more problem.
Some kind soul packaged it for Debian, so it's easy to get if you need it.
- ctecte 5y agoOh nice! Yeah this is basically the same design/intuition that this project also started out with (I also just wanted some way to combine the benefits of aria2c with piping). Wish I found this earlier lol. I only later decided to look into also parallelizing the tar extraction as well. This originally started with looking into making a change to the GNU tar source code itself. However upon cloning I realized there were a lot of design choices in tar that made the assumption of single threaded execution (global state, etc). At this point I decided it would just be easier to re-implement the tar extraction myself using golang's "archive/tar" package.