3 ms·
Some good points. Some incredible engineering has gone into Spark to work around the fact that it runs on the JVM. Memory overhead of Spark particularly (not j
by andygrove 7y ago
Some good points. Some incredible engineering has gone into Spark to work around the fact that it runs on the JVM.
Memory overhead of Spark particularly (not just JVM) is very high. In some cases close to 100x more memory than equivalent query execution with DataFusion [0]. Also you might be interested to see my original blog post with some of my thoughts on this [1].
[0] https://andygrove.io/2019/04/datafusion-0.13.0-benchmarks/ https://andygrove.io/2019/04/datafusion-0.13.0-benchmarks/
[1] https://andygrove.io/2018/01/rust-is-for-big-data/ https://andygrove.io/2018/01/rust-is-for-big-data/
- cozos 7y agoInsightful blog posts! IMO a better memory comparison would be between a Spark executor and a DataFusion ... container I guess (i.e. graphing query time vs spark.executor.memory). This would give you a better idea of memory TCO on a cluster.