4 ms·
Well here’s the thing though, I don’t feel like parallelism is just about speed. I for one like the idea of being able doing 10 things at once, and all tasks r
by devxpy 8y ago
Well here’s the thing though, I don’t feel like parallelism is just about speed. I for one like the idea of being able doing 10 things at once, and all tasks running at the same speed. It’s hard to do with current software, yes. But there a lot of room for improvement, no?
- foldr 8y agoIf you make the processor 10x faster, then you can effectively do 10 things at once without having multiple cores.
- ben-schaaf 8y agoThis isn't quite true. There's a lot more to CPU performance than how fast it can process instructions. Someone configured Intel's 9900k at 500mhz 8 cores and compared it to 4Ghz 1 core. The actual performance characteristics are vastly different. If I had to choose between 10x more cores or 10x faster cores, I'd choose more cores.
- wtracy 8y agoI'm having trouble imagining a performance benefit to two (otherwise identical) 1ghz cores over one 2ghz core, unless: 1. The two processes cause frequent evictions of the CPU cache in some way that two cores can mitigate by having separate caches. 2. You are able to pin a process to a CPU core, and eliminate all the overhead of running the process scheduler. I would be curious to see the benchmark you're referring to. My gut tells me that the performance differences come from comparing different processor families.
- ben-schaaf 8y agoI think this is the source: https://www.youtube.com/watch?v=feO54CBUJCk https://www.youtube.com/watch?v=feO54CBUJCk The whole point of this test is that it's the same processor just configured differently, so there's no architectural or even silicon differences.
- MaulingMonkey 8y agoThe faster your core, the more sensitive you are to latency, even if you have infinite throughput. Say your DRAM has a latency of 65ns - on a single 2ghz core, an L3 cache miss is going to be "twice" as expensive in terms of clock cycles (130) as it would be on a 1ghz core (65). So, to take a super contrived example, if you have a parallel program that needs to run 1B cycles worth of instructions with 1M worth of L3 cache misses. On a single 2ghz core that might take: (2B cycles + 130 cycles / cache miss * 1M cache misses) = 2130 M cycles (1.065 seconds @ 2ghz) On a dual 1ghz core, both sharing a single L3 cache of the same size (so same number of L3 cache misses), that might take: (2B cycles + 65 cycles / cache miss * 1M cache misses) = 2065 M cycles (1.0325 seconds @ 2x1ghz) Which is slightly faster. As long as you're bound by latency rather than throughput, and can perfectly parallelize the problem, more cores instead of more raw clock speed will win out, even with identical hardware and the two cores actually sharing some stuff (the L3 cache) and ignoring the bonuses of fewer L1/L2 cache misses (because they aren't shared, "doubling" your L1/L2 cache.) The machine I'm typing this from has 4MB of L3 cache and I frequently deal with hundreds of gigs of I/O. Suffice it to say, there are a lot of L3 cache misses. Moving away from identical hardware - slower cores can get away with less speculative execution / shorter instruction pipelines, which make things like branch mispredictions cheaper in terms of cycle counts as well. This is why modern GPUs end up with hundreds of cores getting up into the 1Ghz range or so. They deal with embarrassingly parallel workloads, and can get better performance by upping core counts than they can by upping clock speeds.
- wtracy 8y agoIt seems to me that the speed increase in your example comes from having twice the cache and therefore half the misses. You can do that without adding cores. (Barring edge cases where two threads use different pages that happen to share the same cache line.) As for reducing the need for speculative execution: You're just replacing implicit parallelism with explicit parallelism. Whether or not that is a win depends on the workload and the competency of your developers. :-)
- MaulingMonkey 8y ago> It seems to me that the speed increase in your example comes from having twice the cache and therefore half the misses. My math assumed a share L3 cache of the same size, no "twice the cache", no "half the misses". Same number of misses, but each miss is effectively half as costly, because it's only stalling half your processing power (one of your two 1 Ghz cores) for X number of nanoseconds instead of all of it (your single 2 Ghz core). > As for reducing the need for speculative execution: You're just replacing implicit parallelism with explicit parallelism. Yes and no. If you're already explicitly parallel (more and more common), you're just eliminating a redundant implicit mechanism that's based on frequently wrong branch prediction heuristics. One can optimize for those, but then one can argue just how "implicit" it really is... > Whether or not that is a win depends on the workload and the competency of your developers. :-) Of course. But there's a lot of embarrassingly parallel work out there, and improving building blocks, that don't take a geniuses to operate. Pixel shaders, farming video frames out to render farms, compiling separate C++ TUs... these are all things already made explicitly parallel on my behalf. Really maxing out the performance of a single modern ~4 GHz core is no simple feat either.
- gpderetta 8y agoFor 99% of the problems, a 10x faster core will always beat 10x cores. The trade-off is of course different: for the same transistor count/power envelope you can get more than 10x cores.
- devxpy 8y agoIn that case, an important question might be, whether we can actually make it 10x faster. Given a point in the Moore's law curve, can we actually make a single core that's as fast as 10 separate cores? If yes, Will that be more power, space and cost effective?