5 ms·
a lot of people using spark?
by eggie5 9y ago
a lot of people using spark?
- sandGorgon 9y agosame question that i have. Anyone using pyspark in production ? Would you use pyspark mllib in a webservice instead of scikit ?
- wenc 9y ago1) Yes, PySpark is great if you're mostly just doing dataframe manipulation in Spark, using built-in functions. PySpark actually has similar performance to Scala Spark for dataframes. (We've moved away from RDDs) However, if you use a lot of UDFs where Spark has to serialize your Python functions, you might consider rewriting those UDFs in a JVM language. Serialization overhead is still fairly substantial. Arrow is trying to address this by implementing a common in-memory format, but it's still early days. I would still recommend PySpark to most people. It's more than good/fast enough for most data munging tasks. Scala does buy you two things: type safety and low serialization overhead (i.e. significant!), which can be critical in some situations, but not all. Also, the Python way has always been to prototype fast, profile, and rewrite bottlenecks in a faster language, and PySpark conforms to that pattern. 2) Spark MLLib is still fairly rudimentary in its coverage of major ML algorithms, and Spark's linear algebra support, while serviceable, is currently not very sophisticated. There are a few functions that are useful in the data prep stage (encoding, tokenizers, etc.) but overall, we don't really use MLlib very much. Companies that have simple needs (e.g. a simple recommender) and that don't have a lot of in-house expertise, might use MLlib though -- I believe someone from a startup said that they did at a recent meetup. Most of us need better algorithmic coverage and Scikit's coverage is currently much better, plus it is more mature. We also have Numpy at our disposal, which lets us do matrix-vector manipulation easily. There is some serialization cost, but we can usually just throw cloud computational power at it. Also note that for most workloads, the majority of the cost is incurred in training. For models in production, one is typically processing a much smaller amount of data using a trained model, so less horsepower is required.
- sandGorgon 9y agoHi, Thanks for the answer. What you said resonates with me - with a few changes. Spark 2.3 will come with Arrow UDF, that should be a significant performance boost. In that way, yes - we are taking at a forward looking bet. About mllib - yes, we concur with you on algorithmic coverage. And yes, training is the major issue. For example, what I read of Uber's Michaelangelo infrastructure - it seems they train using Spark and save to a custom format that is deserialized (using custom code) and made available as a docker image . There is value in consistency - using Spark thtoy2and through. Wonder what you thought of that ?
- wenc 9y ago1) I've heard about vectorized Python UDFs in Spark 2.3. Thanks for reminding of that. https://databricks.com/blog/2017/10/30/introducing-vectorized-udfs-for-pyspark.html https://databricks.com/blog/2017/10/30/introducing-vectorize... 2) I'm not that familiar with what Uber is doing. My take is I'd like to use Spark for as much as I can, but there are parts that are either more performant or easier to accomplish in Python. Spark with Arrow will definitely change the game.
- threeseed 9y agoAbsolutely Every single large scale data science team e.g. Google, Spotify, AirBnb will be using Spark for most of their work. It is by far the defacto standard for working with large datasets. Especially since it integrates so well with machine learning (H2O) and different languages (Scala, Python, R).
- aretaic 9y agoWe use spark for most of our work. We love it so far, able to handle all our use cases so far and we really appreciate the fact that Scala also runs on the JVM.
- ivanceras 9y agoMaybe too soon, but this framework[0] claims to be 2x faster than spark [0]: https://datafusion.rs/ https://datafusion.rs/
- rajman187 9y agoDefinitely. It's very nice to do large jobs in such a scalable manner. And interacting with databases is very straightforward. I'd also recommend Scala especially if using Spark. I've grown to like it as much if not more than python and you can use zeppelin/jupyter notebooks with is as well.