3 ms·
The 'store' portion buys you a lot of raw performance because its modern architecture friendly. tl;dr, frame it as the row store being an array of structures wh
by Everlag 10y ago
The 'store' portion buys you a lot of raw performance because its modern architecture friendly. tl;dr, frame it as the row store being an array of structures while the column store is a structure of arrays.
Then, your basic way of running a query is a sequential scan over columns. That means queries are going to be more cache friendly while also spending much more time in a trace.
If you sort your data into so-called projections, then the columns you sorted by become the equivalent of sorted arrays, which are extremely compressible. At that point, you can cheaply run-length or delta-encode the sorted columns to achieve incredible compression ratios. Even further, any query using those sorted columns can multiply the effective storage bandwidth available to it by the compression ratio.
I'm not a researcher, just an interested developer, so please feel free to yell at me if anything seems wrong :)
EDIT: Turns out I was actually working on an unpushed branch, its been merged to master. The repo now actually has some interesting work.