4 ms·
So, something between parquet/ORC (a bit more compressed) and arrow (very CPU scan-friendly) I wish we could specify what we intend to do to data, craft a cost
by BenoitP 3y ago
So, something between parquet/ORC (a bit more compressed) and arrow (very CPU scan-friendly)
I wish we could specify what we intend to do to data, craft a cost model, and let a 4th gen system optimize around that. Something that would pick and choose between the different compression techniques in this paper and also the ones from arrow and parquet.
- abeppu 3y agoWhile I wish there were good principled methods for determining these choices, isn't a meaningful limitation that often when you first start writing the data, you typically don't know all of the ways it will eventually be used? Sometimes, the team tasked with making data available for various forms of bulk processing has to guess at what teams and mandates might exist years from when they are setting up the data lake, or even what compliance requirements will arise. E.g. as different areas introduce data protection laws, perhaps your datalake which could previously assume that records are written but not deleted has to support queries to find PII values associated with a person who has issued a delete request, and you're forced to rewrite lots of files. Or a downstream team starts using vector DBs and wants to assemble datasets by doing queries for records which match any of 100k IDs, which have associated vectors in regions of interest, and you need something that's not a single scan but also not a large number of point lookups. Etc, etc. How do you have a principled method for optimizing for unknown future use cases? Does it make sense to talk about a probability distribution over future queries?
- hodgesrm 3y agoThis sounds a lot like Codd's framing of the relational model "problem" in his 1970 ACM paper. [0] His solution was to decouple logical from physical access. That framing and the solution still seem correct to me for the data lake use case. This implies that any lower level storage formats we pick can morph to better optimized forms in a way that does not alter query semantics. It's not possible to solve this without adding refinements to the framing, such as whether data is mutable/immutable, and adding infrastucture to the solution, like metadata management and alternate projections of data. Vertica/C-Store introduced the latter a couple of decades ago. [1] Plus ça change... [0] https://www.seas.upenn.edu/~zives/03f/cis550/codd.pdf https://www.seas.upenn.edu/~zives/03f/cis550/codd.pdf [1] https://web.stanford.edu/class/cs345d-01/rl/cstore.pdf https://web.stanford.edu/class/cs345d-01/rl/cstore.pdf
- hinkley 3y agoAs with most frustrations in my career, many of the expensive bits come down to hoarding. We don’t know what will spark joy so the system has to be able to do anything at any time. This is not free. Sometimes it’s goddamned expensive. I have in this decade encountered systems that still have to run overnight. Meanwhile I’m spending multiple developer salaries maintaining a system that might be asked to answer a question in ten seconds, or might not be asked any questions for days at a time. Somebody save me.
- thesz 3y agoWe, probably, will be even better off by leaving cost model craft to the system. I recently had to look into various TPC benchmarks and some of them are very non-trivial to cost-estimate. I found at several queries in TPC-DS that join the same table to itself four (4) times. Even triangles (join with itself three times) are hard, squares like these in TPC-DS are even harder.
- datadeft 3y agoI would be happy if Parquet or ORC was used in most DWHs. The difference between JSON and Parquet is much bigger than the difference between Parquet and BtrBlocks.