6 ms·
Counterpoint: The NVIDIA Titan RTX has 4608 CUDA cores. We won't see 1024-core CPUs because GPUs cover that use case already.
by timerol 7y ago
Counterpoint: The NVIDIA Titan RTX has 4608 CUDA cores. We won't see 1024-core CPUs because GPUs cover that use case already.
- tombert 7y agoI was going to say that too; GPUs aren't exactly general-purpose CPUs like an Intel/AMD/ARM chip would be, but they clearly prove that having thousands of cores can help specialized cases.
- pmoriarty 7y agoPlease forgive my ignorance, but are GPUs not suitable for general purpose computing because they don't have enough types of logic gates to allow turing complete computation on them, or are they capable of such computation but just aren't optimized for it or what? Is it something software can do anything about?
- Symmetry 7y agoA big GPU might have 50 units capable of independently issuing instructions and each of those might have 50 execution units that will execute that instruction if its program counter matches their own. So it's hard to make a good comparison in thread/core count between modern GPUs and CPUs. But each of those 50 blocks can certainly emulate a CPU core on its own perfectly well. But that would be pretty inefficient compared to just using a CPU core.
- diabeetusman 7y agoTo directly answer your question, graphics cards are Turing complete but they are optimized for a highly parallel workload.
- lalaithion 7y agoGPUs are turing complete. They and CPUs are turing-equivalent, but they're optimized for different tasks. Specifically, GPUs are great at doing the same task on different data all at once.
- wmf 7y agoThis gets back to the example in the blog post. GPUs can run anything, but if the code spends any appreciable time communicating between threads the performance of a GPU will tank because the individual "cores" are quite slow.
- zozbot234 7y agoI don't think it's that simple - you could easily communicate across "threads" in a warp, since these are just SIMD execution units. And you could use GPU main memory to communicate across true "streaming" processors, with some minimum latency. The troubling case is when communication latency limits performance, and the communication is not tightly-coupled within something like a warp. That's when CPU's (or even more exotic models, like systolic arrays etc.) might be preferable.
- lostmsu 7y agoI think the most important part is not communication, but control flow dependencies on data.
- vmchale 7y agoGPUs are SIMD
- petermcneeley 7y agoWhich means given a wave of 64 that gpu has only 72 cores(cus)
- lmeyerov 7y agoyes and no. they're a bunch of distinct simt cores. so what works badly on a few x8s SIMD cores can work great on GPUs. ex: df.groupby('x').agg(my_weird_symbolic_func) can do well in rapids.ai and badly on AVX.
- eesmith 7y agoTo be fair, the original title used here on HN was neither the title of the linked-to essay nor a quote from the essay, but likely a user-submitted title. It has since been changed to match the actual title. Note the first line of the text specifically references "general computing."
- paulddraper 7y agoExactly. If you're looking for advancement in computing power in the same programming paradigms and hardware styles you've always used, you will eventually be disappointed. We've had GPUs for ages. Programming for GPUs is common, though I would suspect has significant innovation remaining.
- jcranmer 7y agoIf you're treating that number as accurate, then I'm running on a 896-core x86 computer. As a comparison between GPUs and CPUs, each "CUDA core" is approximately equal to a SIMD lane on an x86 core. So 56 cores, times 16 lanes (AVX-512) per core, is 896 cores.
- mattnewport 7y ago> each "CUDA core" is approximately equal to a SIMD lane on an x86 core It's not really, that's why Larrabee failed.
- lostmsu 7y agoCould you ELI the difference.
- vardump 7y agoBingo. It has bothered me a long time how the whole industry has fallen to GPU manufacturer marketing speak. That Nvidia Titan RTX thingie is a monster, but it has only 72 SMs (Streaming Multiprocessor), the closest equivalent to a CPU core. That is, SM is the smallest unit that can individually execute a true branch. Just like a CPU core. Not counting hyperthreads here, GPUs have a metric ton of them, because they have to do useful work while waiting for very high latency memory. CPUs need much less of them to cover "bubbles in the pipeline".
- mattnewport 7y agoThe Turing architecture is a bit more sophisticated than that: "Like Volta, the Turing SM is partitioned into 4 sub-cores (or processing blocks) with each sub-core having a single warp scheduler and dispatch unit" "Like we saw in Volta, these changes go hand-in-hand with the new scheduling/execution model with independent thread scheduling that Turing also has, though differences were not disclosed at this time. Rather than per-warp like Pascal, Volta and Turing have per-thread scheduling resources, with a program counter and stack per-thread to track thread state, as well as a convergence optimizer to intelligently group active same-warp threads together into SIMT units. So all threads are equally concurrent, regardless of warp, and can yield and reconverge." GPU 'cores' are still not as general as a full CPU core but they are a fair bit more flexible and sophisticated than a SIMD lane. https://www.anandtech.com/show/13282/nvidia-turing-architecture-deep-dive/4 https://www.anandtech.com/show/13282/nvidia-turing-architect...