4 ms·
The author should spend some time delivering solutions to "real problems". It's easy to be snarky while writing a blogpost with toy datasets but it's difficult
by aub3bhat 9y ago
The author should spend some time delivering solutions to "real problems". It's easy to be snarky while writing a blogpost with toy datasets but it's difficult to actually deliver production solution in an organization. The real COST is often not the time consumed but costs that arise from maintaining heterogeneous architecture and remaining legally compliant.
A database or hadoop cluster is not just an execution engine, but rather a managable, reproducible execution engine. Sure you can code up a smarter algorithm on your personal i9 workstation. However having something available as part of DB or Hadoop ecosystem enables reproducible execution while taking care of compliance (e.g. privacy), security, access control etc.
Eventually we will have a containerized extensions to DB query execution flow that will allow us to have best of both worlds but for 99% practicing data scientists downloading a Gigabyte sized subset to their own workstation is not a viable option.
The real opportunity lies in ability to add arbitrary compute/memory capacity to DB query execution flows.
- zie 9y agoI agree, but at what cost? 10% performance hit, sure, no problem, VM's take about that much from running on real hardware. but a 100% performance hit or worse... you must be crazy. Also, downloading 1GB is no big deal, downloading 1TB is a big deal, so I think your example has more merit in the > 1TB range. i.e. things that don't fit on a laptop easily.
- aub3bhat 9y ago1Mb vs 1Gb vs 1Tb distinction is a function of information contained in data. If its PII even 1Mb is an orginizational headache from security/risk-management perspective. Also most commodity compute nodes have allocation around 16 Gb. The big difference comes in 16~256 Gb range where having a single powerful node can make a huge difference.
- zie 9y agoSure, but 'laptop' can be changed very easily into any container or VM box somewhere that is tied to security/etc but still on a single box. Again sizing here is sort of beside the point, the point is, if its 500% worse to use than just doing it on a single instance/compute node, you better be offering A LOT of magic awesomeness that makes the performance drain acceptable. 10% cost for the benefits that this stuff can give you, that is pretty easy to swallow. 100%-500% perf. cost is just a very, very, very hard to swallow pill.
- JoachimSchipper 9y agoAuditing is a real argument; but is decently-written specialized code really that much less reproducible than Hadoop/SQL? (I'll accept "yes", but note that that also rules out e.g. the fairly common practice of using R or pandas for some ad-hoc processing.)
- aub3bhat 9y agoalso rules out e.g. the fairly common practice of using R or pandas for some ad-hoc processing. This is essentially the reason why all organizations are adopting Spark, since it allows you to write imperative code on dataframes and build ML models.
- rspeer 9y agoNo, "all organizations" are not adopting Spark. You do not need Spark to use dataframes. And distributed computing is terrible for machine learning. Maybe you've worked at a job or two where nobody can comprehend not using distributed computing, as you describe, but it's nonsense to claim that "all organizations" work that way.
- aub3bhat 9y agoNo, "all organizations" are not adopting Spark. All organizations which already have a Hadoop cluster. You do not need Spark to use dataframes. Never claimed this. To clarify Spark allows you to directly port Pandas code while leveraging existing Hadoop cluster infrastructure. And distributed computing is terrible for machine learning. Distributed computing (Both traditional hadoop/spark and latest TF/PyTorch with parameter server) are essential for scaling ML beyond a certain point. Maybe you've worked at a job or two where nobody can comprehend not using distributed computing, as you describe, but it's nonsense to claim that "all organizations" work that way. If you have experience routinely training models on Terabytes of data intended for production deployment. I am happy to hear. There is a vast difference between training a model on your machine for research and building a reliable ML system that scales across large datasets and teams while taking infrastructure costs into account.
- jhayward 9y agoI will note that his local computation could just as easily have been done on a single, fully managed/secured "cloud" node with much the same performance. I think it misses the point to say that not downloading the data is where the cost lies.
- aub3bhat 9y agoI agree but sadly we don't have any plug-in solutions that operate in similar manner. How do you allocate capacity? What if your node has hard upper limit? How do you ensure security e.g. secomp, jails etc. The moment data crosses the DB boundry you need significant investment in ensuring compliance. The reason Spark ecosystem has been so popular is because it enables these types of computations without breaking the model.
- Terretta 9y agoThis line of discussion worries me, in case others think your security assertions are true and could be lead astray. Today’s OSS big data query execution environments are BnL neither secure nor compliant. This is not a knock. They were simply not designed for those constraints. They were built for academic or single tenant use cases without separation of duties / control. This also applies to commercial products such as Splunk. There is tremendous investment going on the last couple years in raising the security and compliance bar. We’ve worked on this with the usual suspects. But barring a handful of proprietary stacks (the big three CSPs, and a couple enterprise on prem bare metal distros) getting close, we are not there yet. Trying to land with your security or risk teams or regulators that the CSP pulled it off will likely take you longer than provisioning a compliant laptop you’d keep locked up with two keys. Unless you have several tens of man years invested in in-depth security wrapping these environments, or can choose a big three CSP w/o answering to anyone, today I’d still recommend the trusted laptop build approach for truly sensitive algorithms and computations.
- aub3bhat 9y agotoday I’d still recommend the trusted laptop build approach for truly sensitive algorithms and computations. You are utterly wrong. Algorithms and computations (especially ML kind) are never sensitive, its the data which is always sensitive. And that ALREADY exists on the cluster. If an Organization already has a Hadoop cluster containing data. You are suggesting that somehow having it downloaded to a secured laptop is better? Than say Spark running on top of the cluster? I think you are deeply mistaken. The cluster instances are already protected (if not you have a bigger problems). Also while an organization might not have Hadoop, they surely have an RDBMS, in which case the algorithms are even more useful.
- geofft 9y agoThat's absolutely true, but I think the author's argument is likely to be that the popularity of Spark and Hadoop in the first place is misguided, and in a world where people were paying attention to performance instead of novelty, we'd have something like Spark and something like Hadoop that was optimized for a single computer, not for scaling for the fun of scaling.
- aub3bhat 9y agoscaling for the fun of scaling. I think your arguments are misguided, for every compute bound task where Hadoop/Spark undeperform, 1000 other ETL type tasks where hadoop is indispensable. As a result any organization running a large hadoop cluster will already have underused compute capacity for free and a maintainance staff which already taking care of the cluster. Thus from an organizations perspective the time difference is not material especially for batch jobs, this is the reason why Presto and Spark have been so successful. They enabled underutilized hadoop clusters to be used for ML and data science while delivering reasonable performance for Zero cost.
- Terretta 9y ago> ”... Staff already taking care of the cluster.” “The” cluster, singular ... This is the catch. If you need security and compliance, you can’t today derive the benefits of re-use by other teams. Given today’s distros, and assuming your threat model needs to account for insider threat, you need a different cluster for each data ownership grouping and data sensitivity level. For four teams with three levels of data, you’d need twelve completely independent clusters. Unless, as noted above, you’ve done a ton of in-house multi-tenancy work to provide full stack security and compliance assurances and audits.
- aub3bhat 9y agoHowever you may split clusters by role/access. You will have I/O bound ETL/Interactive tasks as primary use case. Thus the investment in security will be constant whether you reuse the same infrastructure for compute/memory bound ML tasks or not.
- mcguire 9y ago"A factor of 10x isn't a horrible tax to pay for re-using re-usable infrastructure, like your database. A factor of 60x-130x is more debatable, but I think the CIDR paper is totally worth talking about. I'm a bit disappointed that there aren't reasonable baselines, but not as much as that the SIGMOD paper doesn't have any baselines at all." His "smarter algorithms" aren't.
- lambda 9y ago> However having something available as part of DB or Hadoop ecosystem enables reproducible execution while taking care of compliance (e.g. privacy), security, access control etc. You are conflating a couple of things here. Of course for real analyses, you want to run them on standardized environments with appropriate access controls. But nothing says that "standardized environment" has to use a distributed database that, from all accounts, seems to harm more than help. You could just have a RDBMS or even filesystem, which is appropriately backed up or replicated, on a single beefy VM or physical machine. > The real opportunity lies in ability to add arbitrary compute/memory capacity to DB query execution flows. Only if the tools used to provide that ability actually allow you to take advantage of it. If scaling up to 160 nodes doesn't allow you to be any faster than 1 single laptop, you probably could have spent a lot less money on just adding more RAM to a single beefy server than a 160 node cluster.