3 ms·
I have a question. 125.4 Pflop/s ... that would be about 23k nvidia GPUs (granted, 1080s). They claim they'll be able to do that with 10k by the end of the year
by iofj 10y ago
I have a question. 125.4 Pflop/s ... that would be about 23k nvidia GPUs (granted, 1080s). They claim they'll be able to do that with 10k by the end of the year, beginning of next year with server-class GPUs.
So that Chinese number seems awfully low to me. I would expect that the number Amazon has to be higher, for instance. Same for Microsoft and Google. Therefore I'd be amazed if the DoD, NSA and even the DoE wouldn't have more capacity available.
- Etheryte 10y agoThis isn't about total computing power, but about computing power per one system.
- jobigoud 10y agoBut if you have an application that can be distributed transparently to thousands of GPUs, the difference between one and several systems is not very relevant.
- zhte415 10y agoIt is, because some problems can be distributed to 1000s of GPUs and done in parallel, and some can't, because the answer to one calculation depends on the answer to another calculation. You could combine the computing power of Azure, AWS and Google, and be pretty disappointed because of all the waiting time due to latency from one data center to another That's when the system's architecture - software and hardware - becomes important.
- dekhn 10y agoMany problems are very latency tolerant. Unless you have an algorithm which is truly latency intolerant, I argue you are best served not investing in low-latency interconnect because it costs so much. When I ran Exacycle, we distributed protein folding, protein design, drug discovery, telescope design, and other problems globally. We never had an issue with latency, because these problems all partition really well. People who claim supercomputers are "necessary" for these problems typically construct problems that are well-matched to supercomputers (for example, running molecular dynamics on huge proteins) but they tend not to have very high scientific value. In my experience, partitioning to minimize communication has always increased my total scientific throughput, while programming to supercomputers has always reduced it.
- svantana 10y agoYour arithmetic is in the ballpark -- the article says this machine has 40k processors. But, firstly, these are double precision numbers, I believe that lowers nvidia performance substantially. Furthermore, scaling is not as trivial as putting them in the same room -- you need massive networking, cooling and power supply. Also, while the LAPACK benchmark can work pretty well on GPUs, the stuff that the computers actually is used for might not.
- StreamBright 10y agoIn practice you have many more challenging question that calculate how many GPUs would do the job. For starting how would you connect the GPUs? And so on.