4 ms·
I'm slightly concerned about the Hadoop/Spark + machine learning ecosystem right now. There seem to be a lot of technologies and projects being built up on the
by probdist 11y ago
I'm slightly concerned about the Hadoop/Spark + machine learning ecosystem right now.
There seem to be a lot of technologies and projects being built up on the core foundation of in memory distributed computation. MLlib in Spark, H2O, Apache Mahout/Samsara, probably many others.
My impression is Samsara and SystemML's DML are being designed to supply the primitives you need to build a machine learning model with less awareness of the underlying system model. So this is ostensibly to save the developer of a new algorithm the headache of thinking how their algorithm needs to relate to the distributed computation model like one should consider for developing with H2O or Spark. However, It seems like many common off the shelf algos are already available and implemented indicating that perhaps working with Spark or H2O backends directly is not that hard to produce performant ML codes.
As a computational scientist learning a DSL to accelerate development of an algorithm of mine into production via Spark seems like a potential distraction from just getting better with Spark.
- urlwolf 11y agoThe current Big data ecosystem is a difficult place to be for companies that bet the house on hadoop (cloudera, Hortonworks, MapR) for sure. There are new technologies coming at a breakneck pace. Check out apache Flink if you need streams and the microbatches in Spark don't do it for you. The ML lib there is not that far ahead yet though.