3 ms·
As of a few days ago (Apache Spark 3.2 release), you can use the pandas API on a Spark cluster: https://databricks.com/blog/2021/10/04/pandas-api-on-upcoming-ap
by jointpdf 5y ago
As of a few days ago (Apache Spark 3.2 release), you can use the pandas API on a Spark cluster: https://databricks.com/blog/2021/10/04/pandas-api-on-upcoming-apache-spark-3-2.html https://databricks.com/blog/2021/10/04/pandas-api-on-upcomin...
- alextheparrot 5y agoHow the relevant lesson from the original post was “Actually, you can scale pandas these days” eludes me.
- dikei 5y agoSince the original post criticizes the behavior of Panda runtime where it loads everything into memory before processing, I think replacing the underlying Panda runtime with Spark runtime is very much a valid solution. You can keep using Panda code for exploratory work, then with a little effort move it to production with Spark runtime.
- monkeybutton 5y agoAs an avid user of pandas: That solution is just kicking the can down the road.
- ldng 5y agoAnd now you have a whole set of new problems ?