4 ms·
scaling for the fun of scaling. I think your arguments are misguided, for every compute bound task where Hadoop/Spark undeperform, 1000 other ETL type tasks w
by aub3bhat 9y ago
scaling for the fun of scaling.
I think your arguments are misguided, for every compute bound task where Hadoop/Spark undeperform, 1000 other ETL type tasks where hadoop is indispensable. As a result any organization running a large hadoop cluster will already have underused compute capacity for free and a maintainance staff which already taking care of the cluster. Thus from an organizations perspective the time difference is not material especially for batch jobs, this is the reason why Presto and Spark have been so successful. They enabled underutilized hadoop clusters to be used for ML and data science while delivering reasonable performance for Zero cost.
- Terretta 9y ago> ”... Staff already taking care of the cluster.” “The” cluster, singular ... This is the catch. If you need security and compliance, you can’t today derive the benefits of re-use by other teams. Given today’s distros, and assuming your threat model needs to account for insider threat, you need a different cluster for each data ownership grouping and data sensitivity level. For four teams with three levels of data, you’d need twelve completely independent clusters. Unless, as noted above, you’ve done a ton of in-house multi-tenancy work to provide full stack security and compliance assurances and audits.
- aub3bhat 9y agoHowever you may split clusters by role/access. You will have I/O bound ETL/Interactive tasks as primary use case. Thus the investment in security will be constant whether you reuse the same infrastructure for compute/memory bound ML tasks or not.