3 ms·
Not only that, - How do you make sure the users are not writing badly optimized tables/pipelines that end up consuming too many resources. - How do you facili
by crorella 6y ago
Not only that,
- How do you make sure the users are not writing badly optimized tables/pipelines that end up consuming too many resources.
- How do you facilitate data discoverability so they don't end up creating a new table where 90%+ of the data is already present in another already existing one?
- How do you make sure they are mindful with the way the model the tables such there are not too many files due to bad partition/bucketing, compression is leveraged and good datatypes are picked?
- atomicnumber3 6y agoI work in this space and the questions outlined in you and your parent's comment are absolutely spot-on. The worst part is, these things tend to grow organically where the original ancestor of everything is engineer #3 of the 5-person startup who decided to write a cronjob to dump the prod db every night to feather files on a samba share. Then one cron job becomes three. Then ten. Then you're using S3. Then there's dependencies. Then you're using Luigi/Airflow. Then you're using Spark. Then you're constantly messing with partitioning and performance. Then you're using Hadoop and YARN and configuring queues and capacity scaling. At this point the "infra" team is 10 people and there's 20 data scientists. Then you triple in size a couple dozen times and now it's time to figure out how to retroactively slap security controls on top of everything. That turned into a bit of a rant. Honestly I love this space but I agree, as hard as the software bits are, the hard part is the entire picture as a whole.