4 ms·
I'm surprised the GPU is a win when the data is coming from GCS. The CPU still has to touch all the data, right? Or do you have some mechanism to keep warm data
by Boxxed 1y ago
I'm surprised the GPU is a win when the data is coming from GCS. The CPU still has to touch all the data, right? Or do you have some mechanism to keep warm data live in the GPUs?
- winwang 1y agoYep, CPU has to transfer data because no RDMA setup on GCP lol. But that's like 16-32 GB/s of transfer per GPU (assuming T4/L4 nodes), which is much more than network bandwidth. And we're not even network bound, even if there's no warm data (i.e. for our ETL workloads). However, there is some stuff kept on GPU during actual execution for each Spark task even if they aren't running on the GPU at the moment, which makes handling memory and partition sizes... "fun", haha.
- threeseed 1y agoI've used GPU based Spark SQL for many years now and it sounds flashy but it's not going to make a meaningful difference for most use cases. As you say the issue is that you have an overall process to optimise from getting the data off slow GCS onto the nodes, shuffling it which often then writes it to a slow disk before the real processing even starts then writing back to a slow GCS.
- winwang 1y agoNot sure what your use cases are, but I haven't had too much issue seeing good gains vs bare Spark -- GCS has not been my bottleneck.
- _zoltan_ 1y agowould you be able to share a runtime with operator breakdown for the curious ones among us?
- winwang 1y agoThat's a pretty interesting idea, might take a bit to prepare a useful graphic/post. Also, what do you think would be the best way to structure such a post? But, here's a small bit of something perf-y: during large shuffles, I was able to increase overall job performance/efficiency by using external shuffles, even with times of ~5s median shuffle write for a couple hundred MB partitions (I hope I'm remembering this correctly, lol). This is not particularly great, but it did allow for cost-efficiently chewing through some rather large datasets without dealing with memory issues. There's also an awesome side benefit in that it allows us to use cheap spot workers in more scenarios.