4 ms·
Ok that seems like magic. How is it able to only read a subset of the data like that? Is this because of the parquet file format? Also can the same thing work
by isuckatcoding 3y ago
Ok that seems like magic. How is it able to only read a subset of the data like that? Is this because of the parquet file format?
Also can the same thing work for files hosted on s3?
- simonw 3y agoYes, this works for S3 files as well. The trick it's using is the HTTP Range header, which lets a client request eg bytes 4500-4900 of an HTTP file. Most static file hosting platforms - S3, GCS, nginx, Apache etc - support Range headers. They're most commonly used for streaming video and audio. The Parquet file format is designed with this in mind. You can read metadata at the start of the file and use it to figure out which ranges to fetch. Columns are grouped together, so sum() against a column can be handled by fetching a subset of the file.
- mosselman 3y agoThank you for explaining. I never really looked into this to understand it and because of that it felt like magic, which is always in indicator that you just don’t understand something. I was going to add “in tech”, but it is an indicator of that with anything in life.
- BenoitP 3y agoParquet is columnar, you can read each column independently, and a different compression scheme can be used for each column. You first have to read the beginning of the file to get the layout though (And the end of it as well, according to a library I have been using. Don't know why) . The http server must of course support range queries.
- choppaface 3y agoRight, there's all sorts of metadata and often stats included in any parquet file: https://github.com/apache/parquet-format#file-format https://github.com/apache/parquet-format#file-format The offsets of said metadata are well-defined (i.e. in the footer) so for S3 / blob storage so long as you can efficiently request a range of bytes you can pull the metadata without having to read all the data.