4 ms·
Big data is about I/O not CPU I’m a C++ veteran btw and understand the point but big data is about how to process petabytes of I/O not how to consume CPU.
by jpz 7y ago
Big data is about I/O not CPU
I’m a C++ veteran btw and understand the point but big data is about how to process petabytes of I/O not how to consume CPU.
- sriram_malhar 7y agoYes, you are right. That said, the COST paper's presumption is that not every "big data" problem is big (petabyte size), that most probably fit in a single system's disk & memory.
- emsy 7y agoCorrect me if I’m wrong, but isn’t the point of writing cache coherent code the fact that memory I/O is the bottleneck these days?
- growse 7y ago> Big data is about I/O not CPU > I’m a C++ veteran btw and understand the point but big data is about how to process petabytes of I/O not how to consume CPU. I'm not sure this is cut-and-dry. Back when I was working on Spark workloads, there was some interesting research being done on where the bottlenecks were for jobs. I think it turned out for a lot of jobs, infinite disk / network io didn't give as much of an improvement as you'd expect. https://databricks.com/session/making-sense-of-spark-performance https://databricks.com/session/making-sense-of-spark-perform...
- philipov 7y agoNot for everyone. In finance, a query might only use gigabytes or terabytes of data, but need to do a ton of simulation and calculation on top of that. Optimization of e.g. trading algorithms is entirely CPU-bound.
- jpz 7y agoYes, I'm aware. I've written and maintained several of these systems in exotic derivatives space. Scale-out of CPU was required for these. These are not big data problems, however, by definition.
- bane 7y agoThis is true in a "water is wet" kind of way. The point is that a great many problems that can fit neatly into a single machine are being turned into I/O problems by being distributed onto clusters. There's an incredible number of gigabyte and even terabyte scale problems that are consuming racks of blades when a little thinking and understanding of the problem being solved can be done pretty nicely on far fewer resources. What's really happening is that to many people they think it's "easier" to simply rack more equipment into the cluster and end up shifting the complexity into cluster administration rather than programmer time.
- mannykannot 7y ago>...and end up shifting the complexity into cluster administration rather than programmer time. Sometimes, that is the right thing to do. The problem is when the cluster solution also adds to the programmer-time complexity.
- zozbot234 7y agoEven CPU is about "I/O" these days. Memory (RAM) is the new disk - and memory bandwidth is generally the performance bottleneck in heavy workloads, especially in multicore. This might be one reason why loosely 'C-like' languages like Rust are going back in style. High-level languages are terrible for memory bandwidth.