3 ms·
I've used GPU based Spark SQL for many years now and it sounds flashy but it's not going to make a meaningful difference for most use cases. As you say the iss
by threeseed 1y ago
I've used GPU based Spark SQL for many years now and it sounds flashy but it's not going to make a meaningful difference for most use cases.
As you say the issue is that you have an overall process to optimise from getting the data off slow GCS onto the nodes, shuffling it which often then writes it to a slow disk before the real processing even starts then writing back to a slow GCS.
- winwang 1y agoNot sure what your use cases are, but I haven't had too much issue seeing good gains vs bare Spark -- GCS has not been my bottleneck.
- _zoltan_ 1y agowould you be able to share a runtime with operator breakdown for the curious ones among us?
- winwang 1y agoThat's a pretty interesting idea, might take a bit to prepare a useful graphic/post. Also, what do you think would be the best way to structure such a post? But, here's a small bit of something perf-y: during large shuffles, I was able to increase overall job performance/efficiency by using external shuffles, even with times of ~5s median shuffle write for a couple hundred MB partitions (I hope I'm remembering this correctly, lol). This is not particularly great, but it did allow for cost-efficiently chewing through some rather large datasets without dealing with memory issues. There's also an awesome side benefit in that it allows us to use cheap spot workers in more scenarios.