4 ms·
> At the same time, data keeps getting bigger and computers come with more and more cores (which Python cannot easily take advantage of), while single-core perf
by munro 9y ago
> At the same time, data keeps getting bigger and computers come with more and more cores (which Python cannot easily take advantage of), while single-core performance is only slowly getting better. Thus, Python is a worse and worse solution, performance-wise.
PySpark is makes it really easy to take advantage of multiple cores & machines. Most operations I want to do to my data I can find in PySpark's pyspark.sql.functions, so I get all the benefits of the JVM. In the cases I need something from Python, I can just UDF, it's a little slower than JVM but still extremely fast when distributed--I find all problems come down to time or memory complexity, which is independent to whatever your programming in. Also, it's very easy to take advantage of spot instances with Spark... I'm usually working with 2-20 spot instances, and sometimes go up to 60 depending on what I'm doing.
- pcx 9y agoThis is true, but using a Distributed system like Spark itself adds a ton of complexity in having to understand and manage it. If one can do something with a set of stateless processes, even if it's more performant, I feel it's a bad idea to use a distributed system instead. Not always, but a good majority of cases that I've seen. I've seen projects where Celery would be enough but instead they chose to use Spark/Storm and never delivered.
- deleted 9y ago[deleted]
- quietbritishjim 9y agoThe original article said that one reason it doesn't matter that pure Python's performance is poor is that you can use numpy (and pandas) to vectorise things, which then has native code performance. It goes on to say that his current problem is that the things he's doing today can't be vectorised with numpy – so that poor performance does matter after all. If he can't even express his code in terms of numpy operations (and other C-based libraries like scipy), I doubt they're going to be expressible in terms of Spark's primitives, which are a considerably smaller subset.
- lmm 9y agoSpark is great, but at that point why not just use Scala? It offers Python-like conciseness/productivity, and by using the same language Spark is written in you avoid a big class of possible interop issues.
- cjalmeida 9y agoPySpark should only be used for prototyping. It add an enormous extra overhead on operations due to serializing data back and forth between the Java and Python processes.