3 ms·
He means the other way around: I/O improvements are outpacing CPU improvements. Native refers to (I think) the following issues: https://issues.apache.org/jir
by mydpy 11y ago
He means the other way around: I/O improvements are outpacing CPU improvements.
Native refers to (I think) the following issues:
https://issues.apache.org/jira/browse/SPARK-12785 https://issues.apache.org/jira/browse/SPARK-12785
https://issues.apache.org/jira/browse/SPARK-8641 https://issues.apache.org/jira/browse/SPARK-8641
Code generation is enabled by SPARK-8641, but not sure exactly what it entails. I think it is related to some of the RDD transformation/action merging they do to optimize runtime operations in 2.0.
Your thoughts?
- vvanders 11y agoNot overly familiar with Spark but usually the bottleneck for CPUs is access to DRAM. Unless you're doing completely linear reads or prefetching appropriately(which almost no one does right) you'll be cache-miss bound and your CPUs will be idle waiting for data.
- azth 11y agoThat's what I was thinking. The slide seems to imply that workloads are becoming CPU bound due to IO speeds increasing a lot.
- zzalpha 11y agoThat's precisely what he's saying. This is the thesis of those who have been watching the explosion in solid state disks. The claim is that bulk I/O is becoming so damn fast that the pendulum is swinging toward computing power being the new limiting factor.
- vvanders 11y agoNo expert here on SSDs but it still seems like they're off by an order of magnitude. PCIe SSDs appear to top out around 1.8GB/sec while DDR4 is in the 12-19 GB/sec range. Either way CPU processing power has never really been a constraining factor(since most people don't want to write data aware code), it's always been bus-bound be it DRAM, IO or other.
- lmm 11y agoSpark is built around keeping things in-memory. If there's any workload that can max out a CPU (and surely there must be, otherwise no-one would bother building faster CPUs) then there's probably someone doing it on Spark (linear algebra? monte carlo simulation?). Spark might be a good fit with RDMA (looks like there's some initial work in that direction).
- haimez 11y agoIt entails fusing "narrow" operations on data frames into a single method of generated code to avoid virtual method invocations, help the JIT, and improve hardware branch prediction among other things. Basically you take the general, composable API functions and generate specific equivalent code that avoids the overhead of dealing with abstract interfaces like Iterators.
- vvanders 11y agoNifty. I've always been a big fan of code generation. How does Spark guarantee data memory contiguity(or do they at all)? Do they use misc.sun.unsafe or some form of memory management?
- haimez 11y agoYeah, since spark 1.6 (some optimizations in spark 1.5) the data being operated on is managed "off heap" which means it doesn't add to garbage collection times and is stored in a more contiguous and cache friendly layout. It's definitely using Unsafe but can't speak to the exact implementation.
- azth 11y ago> Basically you take the general, composable API functions and generate specific equivalent code that avoids the overhead of dealing with abstract interfaces like Iterators. That's pretty cool. Similar to what Rust already does when the compiler elides allocations for chained iterators. I wonder if value type support in the JVM would avoid the need to do this type of code gen.