3 ms·
This is true when you compare the performance vs Java/Scala, but if you compare it with other tools that are native in Python, it is not really much worse. For
by rxin 11y ago
This is true when you compare the performance vs Java/Scala, but if you compare it with other tools that are native in Python, it is not really much worse. For examples, Pandas operations that use custom UDFs are substantially slower than the native operations.
That said, as part of Project Tungsten, we have some ideas about a batch columnar format that can be shared by Python, R, Scala and Java, and that should be able to eliminate most of the inefficiency in serialization across process boundaries.
- mziel 11y agoThat sounds very interesting. Is there any ticket, where I can follow the progress on the batch columnar format you mentioned? Btw, I was critical about the issue above, but I do love Spark, using it on a daily basis. :)
- rxin 11y agoI just created a JIRA ticket tracking this: https://issues.apache.org/jira/browse/SPARK-12635 https://issues.apache.org/jira/browse/SPARK-12635 Thanks for the reminder!