8 ms·
Inside the Titan Supercomputer: 299K AMD x86 Cores and 18.6K Nvidia GPU Cores
- hendzen 14y agoThis is pretty awe-inspiring but as a programmer I know it would be fairly difficult to use this machine for existing workloads because so much code would have to be rewritten from typical x86 code to CUDA/OpenCL to use all those GPUs. Personally, I'm more excited for the next wave of supercomputers built with racks of Xeon Phis [1]. [1] - http://www.intel.com/content/www/us/en/high-performance-computing/xeon-phi-for-researchers-infographic.html.html http://www.intel.com/content/www/us/en/high-performance-comp...
- Osmium 14y agoIn fairness, I don't think there are any existing workloads that would benefit 299k cores that aren't already massively parallel :) If you need this kinda thing, your code is already going to be ready for it. I'm with you on the Phis though. Can't wait for one to be affordable for a home machine
- tmurray 14y ago(full disclosure: used to work for NV on CUDA and did very extensive work on Titan, so I am probably biased) If you think your existing MPI app is going to automatically scale to a heterogeneous architecture (high-power x86 on the main CPU, Xeon Phi cores on the accelerator) and get acceptable performance, sorry, it's not going to happen. The fundamental constraints on 2012/2013 Xeon Phi performance that determine how apps should be written are exactly the same as current desktop GPUs (small, high-latency local memory that is not coherent with the rest of the system; relatively slow, high-latency link to CPU; ugly interactions with network cards in most environments; fundamental need to hide memory latency at all times). For any sort of performance beyond a standard Xeon, you're going to want to run a Xeon Phi as a targeted accelerator rather than offloading entire processes to it and using a standard MPI stack. This means you're going to be running in a hybrid host/device mode and using compiler directives or a specific parallel language and API to deal with on-chip execution and data transfer, which puts you in exactly the same solution space as with GPUs. in other words: the Phi of today is not a panacea. you get better tools and more flexibility in terms of the programming model, but the fast path that any of its intended market would use in applications looks identical to GPUs.
- batgaijin 14y agoTo my understanding GPU's basically suck at anything with decision paths/move away from straight matrix manipulation/signals analysis, right?
- deleted 14y ago[deleted]
- deleted 14y ago[deleted]
- confluence 14y agoGPUs suck at any problem that cannot be easily divided. If you can map a function over arbitrary chunks of self-contained data GPUs will perform better.
- vilya 14y agoGPUs are SIMD machines, so they're executing the same instruction simultaneously on all the active cores. That means if you have code which branches, it has to mask out the cores which follow branch B while it executes branch A; then has to mask out all the cores which follow branch A while it executes branch B. In other words, if at least one core follows each side of the branch, it has to execute both branches. If all cores branch in the same direction, you don't get that penalty. A large part of optimising for the GPU comes down to arranging your data and code so that this can happen.
- Cogito 14y agoThe full page (print version) version of the article is at http://www.anandtech.com/print/6421 http://www.anandtech.com/print/6421
- asdfs 14y agoTitle should be "nVidia GPUs", not "nVidia GPU cores".
- paulsutter 14y agoDoes anyone know why they have a separate disk IO system when they could more easily just plug drives into each node/motherboard for higher aggregate throughout, less complexity, and a lower overall cost? EDIT: Blade systems or no, the drives have to physically be placed somewhere. Having a separate subsystem can only take up more space, not less. Two reasons I can think of: (1) independent scaling of compute and storage, and (2) lack of software for a distributed filesystem. Most likely (2) plus inertia is the real reason, all the others seem like rationalizations. For example, they are either able to take nodes offline or they aren't. The need exists whether or not the disks are attached there.
- tmurray 14y agoThey're blade systems, there's not really room for disks. Easier to keep it centralized and easily serviced.
- T-Winsnes 14y agoMy guess is that most of the work they do on this machine won't be bottlenecked on the hdds, so they don't worry about it too much. And it's easier to replace hard drives in a separate rack than taking blades offline to do it.
- caf 14y agoIf you look through the gallery you'll see that their disk subsystem is using a distributed filesystem, Lustre (http://www.lustre.org/ http://www.lustre.org/).
- paulsutter 14y agoThey would need something much more like GFS than Lustre. Lustre is designed for the exact approach they're using.
- willvarfar 14y ago> Does anyone know why they have a separate disk IO system It would be very interesting if they could describe all their general design decisions, such as this > ... when they could more easily just plug drives into each node/motherboard for higher aggregate throughout, less complexity, and a lower overall cost? doesn't the fact that they haven't put the drives on the compute nodes make you question your claim that it would have been 'easier' and 'better' and 'cheaper'?
- caf 14y agoThe hexadecimal numbers in the design on the front panels of the racks appear to say in part: ...Computing Oak Ridge National Laboratory Le... (not too surprising, I suppose ;)
- mtgx 14y agoHow can Anandtech make such a big mistake? It's 46 million GPU cores, not GPU's.
- tjaerv 14y ago177 trillion transistors in total.
- zspade 14y agoMore Transistors than there are synapses in the human brain (I'm aware they are not a direct analogue).