3 ms·
Because most people don’t have an HPC background, aren’t familiar with parquet internals, don’t know how to make their language stream data instead of buffering
by memset 2y ago
Because most people don’t have an HPC background, aren’t familiar with parquet internals, don’t know how to make their language stream data instead of buffering it all in memory, have slow internet connections at home, are running out of disk space on their laptops, and only have 4 GB of ram to work with after Chrome and Slack take up the other 12 GB.
15 GB is a real drag to do anything with. So it’s a real pain when someone says “I’ll just give you 1 TB worth of parquet in S3”, the equivalent of dropping a billion dollars on someone’s doorstep in $1 bills.
- vladsanchez 2y agoFunny analogy! I loved it. I'm ready to start with ScratchData which btw and respectfully never heard of. Thanks again for sharing your tool and insightful knowledge.
- fifilura 2y agoHow do you see the competition from Trino and Athena in your case? Depends a lot on what you want to do with the data of course, but if you want to filter and slice/dice it, my experience is that it is really fast and stable. And if you already have it on s3, the threshold for using it is extremely small.
- fock 2y agowhat is your point? They talked about 15GB of parquet - what does this have to do with 1TB of parquet? Also: How does the tool you sell here solve the problem - the data is already there and can't be processed (15GB - funny that seems to be the scale of YC startups?)? How does a tool to transfer the data into a new database help here?
- wodenokoto 2y ago> How does a tool to transfer the data into a new database help here? Maybe because the problem literally is "how to transfer this data into a database"