3 ms·
I am genuinely curious to see what types of "compute-intensive" applications fit the bill here. Outside of storage workloads (syncing data, etc.), why would yo
by diffserv 7y ago
I am genuinely curious to see what types of "compute-intensive" applications fit the bill here. Outside of storage workloads (syncing data, etc.), why would you need a 100x improvement in data transfer rates between the machines? (We have TPUs and their specialized network architecture for ML-like workloads ...)
Physical distance between the machines in a DC prevents RAM style "shared-memory" architectures, at least ones that aim to have 30~60ns access times (10-20 meters). Unless there are new paradigms for computation in a distributed setting, I don't see the benefit for this ...
Also, what are the fundamental limitations/research problems of todays hardware that prevent us from building a 400G NIC? I cannot think of anything outside of PCI-e bus getting saturated. We already have 400G ports on switches ...
- scottlocklin 7y ago>I am genuinely curious to see what types of "compute-intensive" applications fit the bill here. Well, for example, physical simulations using finite element or boundary value approaches. Pretty much anything you'd use with MPI or do on a supercomputer is going to run better on a machine with a nice network stack like this. Even large scale storage (think backtesting on petabytes of options data) that uses a map-reduce paradigm and is properly sharded for the data access paths and aggregates would benefit from something like this.
- diffserv 7y agoDo you have numbers or papers that support your argument? That these applications are bottlenecked by the network? There is a 2015 paper [1] that argues that improving network performance isn't gonna help MapReduce/data analytics type of jobs much: " .. none of the workloads we studied could improve by a median of more than 2% as a result of optimizing network performance. We did not use especially high bandwidth machines in getting this result: the m2.4xlarge instances we used have a 1Gbps network link." Granted things might have changed by now, but I am curious to see how and by how much? [1]: https://www.usenix.org/system/files/conference/nsdi15/nsdi15-paper-ousterhout.pdf https://www.usenix.org/system/files/conference/nsdi15/nsdi15...
- scottlocklin 7y agoThe 2015 paper is obviously wrong or selling something; it's virtually always IO bound. Yes, I know many people assert otherwise; they're wrong. Some map reduce loads, especially the kind that people running spark clusters want to do, end up moving a lot of data around. Either because the end user isn't thinking about what they're doing (95% of the time they're some DS dweeb who doesn't know how computers work), or because they need to solve a problem they didn't think of when they laid their data down. I guess I cite myself, having done this sort of thing any number of times, and helped write a shardable columnar database engine which deals with such problems. If you don't want to cite me; go ask Art Whitney, Stevan Apter or Dennis Shasha, whose ideas I shamelessly steal. FWIIW around that timeframe I beat a 84 thread spark cluster grinding on parquet files with 1 thread in J (by a factor of approximately 10,000 -the spark job ran for days and never completed), basically because I understand that, no matter how many papers get written, data science problems are still IO bound.
- gnufx 7y agoThere are references somewhere under http://nowlab.cse.ohio-state.edu/ http://nowlab.cse.ohio-state.edu/ for instance.
- gnufx 7y agoI'm not sure about the typical demands on the fabric of FE-type codes, but a typical HPC cluster (university or national) is likely to spend much of its time on materials science work using DFT, which requires low latency rather than high bandwidth for short messages. Somewhere under http://archer.ac.uk http://archer.ac.uk there are usage statistics as an indication of the workload, though it varies month-to-month.