2 ms·
Hi whinvik, we agree that development in Spark is hard, and that is part of the motivation of Fugue. Spark code couples the distributed orchestration and busine
by kvnkho 4y ago
Hi whinvik, we agree that development in Spark is hard, and that is part of the motivation of Fugue. Spark code couples the distributed orchestration and business logic together.
By keeping your code in native Python or Pandas, it will be much easier to develop, debug, and maintain the business logic because your tracebacks will be in native Python. Fugue then takes it to Spark when you are ready to scale.
- whinvik 4y agoI appreciate your response but that is not what I was getting at. I understand that with this you only have to write Pandas and then not worry about scaling. First, I think PySpark syntax is much better than the insanity than is Pandas but if you really like Pandas then you can always use Pandas UDF which Spark supports. But let's say that writing only in Pandas is the preferred way. Now comes the magic part. How do I know that it is using the best join? Will it optimize for spills? Will there be OOM's? These are the things we need to worry about which often lead us needing to go deep inside Spark magic. Now if there's another level of magic which is Pandas to Spark transpiling as I imagine you do here, then I have even less of an idea how to tune it. Again I appreciate you are solving a specific problem in a nice way but I feel like we are actually making the problem even more complicated.
- kvnkho 4y agoAh I understand what you are saying. Not trying to aggressively make you try Fugue out, I appreciate the comments and just want to clarify some stuff. There is no transpiling, we think that is too magical. It's more of routing to the appropriate functions/method. For example, applying the Pandas function per partition. We use Pandas UDF under the hood when it makes sense. Sometimes, we use map partitions as well, it depends on what the user wants to do. So we simplify usage around these methods with minimal overhead (we benchmarked it). Fugue can be adopted as minimally as needed so if you feel the need to write native Spark code or tune Spark, we don't block that from the user. The older Fugue interface was bad in allowing this, so we reworked Fugue to have a suite of standalone functions compatible with Spark/Dask/Ray/Pandas DataFrames. Our opinion (which is totally fine if you disagree), is that most workloads don't really specific features of Spark. There are times when it makes sense to write native Spark code for sure, and for those Fugue won't be a good fit. Our job is to make sure it's not an all or nothing thing that requires the user to make compromises.