4 ms·
> codes that would be effectively computationally intractable if you tried to run them in, say, AWS. such as?
by moab 4y ago
> codes that would be effectively computationally intractable if you tried to run them in, say, AWS.
such as?
- nine_k 4y agoMaybe weather prediction and other problems of fluid dynamics?
- jandrewrogers 4y agoLarge-scale problems that look like computational fluid dynamics, graph analysis, spatiotemporal analysis, various kinds of computational science and physics models, etc. Just about anything that involves modeling and analyzing behaviors in the physical world. All of these involve high-bandwidth parallel orchestration at the data structures and algorithms level. The quality of the network is everything if your code is competently parallel. They don’t run these things on supercomputers for fun, it would be much easier if you could run these codes on AWS. The cloud has many advantages, but high-quality inter-node bandwidth and topology isn’t one of them. In HPC, the network is the most important part of the system.
- moab 4y agoTotally agree with your last point. But based on my understanding, over the past 5-ish years commodity multicores with multi-terabyte memories have made a big dent in the supremacy of supercomputing, at least for some of the topics you mention (thinking about graph analysis in particular).
- jandrewrogers 4y agoGraph analysis was one of my core research areas in my supercomputing days. You can do it on commodity hardware with some caveats. This problem is intrinsically cache-unfriendly, which makes it interesting but also slow even when it fits in memory on a single machine. We made a lot of progress on approaching the theoretical throughput bound on real hardware and even made it work reasonably well when storage-backed, with some caveats. That said, multi-terabyte memories won’t solve interesting problems; we already had that. When I was working on this 15-ish years ago, the real-world data models had trillions of vertices, never mind edges. And that has only gotten larger with time. A lot of the research ended up focusing on the problem of how do you boil the ocean selectively and incrementally to optimize throughput. There is no way to trivially throw hardware at the problem; graph-cutting is hard, and you have to do it even within single servers. Even with sophisticated latency-hiding, it ends up being about effective bandwidth in a context where caches are almost useless. For graph analysis specifically, we could do a lot with big servers, this is true. But it would require a completely different software architecture to the way most graph analysis is done now. This is perpetually on my “copious spare time” lists of projects because there is a big gap here.
- moab 4y agoMakes sense. I hope you keep working on this. There is a big gap between what people are doing in the academic (and even publicly-described industrial literature) and what you describe. So I think there is a lot of opportunity to push on this front if you have some ideas.
- dekhn 4y agoevery time a commodity multicore machine gets better, the supercomputer folks just switch to a larger problem that wouldn't fit on a single machine. Their goal is to engineer a system/build a code that reaches peak performance limited primarily by the physical constraints of the biggest systems.
- jandrewrogers 4y agoTo expand a bit, I did R&D for a supercomputing companies on software latency-hiding many years ago. The idea was that we could run many of these codes on commodity hardware by taking ideas from exotic latency-hiding silicon, which had no direct analogues that could be implemented in software, and inventing something similar for software. This was surprisingly successful! Unfortunately, this had two practical problems. First, it turns out that developers are quite poor at reasoning about control flow in latency-hiding architectures generally. It is analogous to reasoning about very complex lock graphs but worse. Second, you still need prodigious quantities of high-quality network bandwidth and topology, even if latency matters much less, for typical HPC problems. At which point you are half the way to a traditional HPC network anyway. There is still a space for this research in non-HPC applications (like join parallelization in databases), but for HPC the cost-benefit ratio pushes everyone to purpose-built networks. Less learning curve for the devs and you get the exceptional bandwidth and topology the codes need anyway.
- jdkee 4y agoNuclear weapons simulations.
- JonChesterfield 4y agoI think that's the other side of the DoE labs, though it's all pretty opaque to an outsider.