3 ms·
I actually just read a paper where some new file format achieved far faster and smaller writes than the Parquet used in all of these modern big datalakehouses,
by 392 3y ago
I actually just read a paper where some new file format achieved far faster and smaller writes than the Parquet used in all of these modern big datalakehouses, and it was done by guessing like 20 different simple compressions on a sample to figure out what will work best for each chunk of data. Sounds like manual feature detection to me!