4 ms·
The Problem is not with C alone however. Speculative Execution, branch prediction and look ahead to the next 25 instructions, wasting huge amounts of power,...
by usrbinbash 5y ago
The Problem is not with C alone however.
Speculative Execution, branch prediction and look ahead to the next 25 instructions, wasting huge amounts of power,...I mean, what?
The article even mentions that there is another way:
> In contrast, GPUs achieve very high performance without any of this logic, at the expense of requiring explicitly parallel programs.
Yes, C doesn't support that very well. But it could, and other languages, namely Rust, Go & Julia already can. So maybe it's time to do in CPU design what Go did in language design, and hit the brakes on complexity?
We don't need smarter processors, we need processors that can do high throughput on many execution units, and languages that support that well.
- jcelerier 5y agoPlease no :( unparallelizable tasks are already slow enough. My guitar effects rack (where effects all have to be processed one after each other, by design) uses barely less CPU% in a 2020 computer than in a 2011 one and it is extremely frustrating.
- usrbinbash 5y ago> uses barely less CPU% in a 2020 computer than in a 2011 one and it is extremely frustrating. But if that's the case, what's the point in all that extra-complexity in the CPU, if in the end the benefits seem to be miniscule?
- jcelerier 5y agoall this "extra complexity", branch prediction, pipelines, multiple levels of cache, speculative execution was mostly there since the late 80s, 90s in CPU design ; the Pentium pro already had all of this to some level. The last decade was in large part about more SIMD and more cores and it's been a real PITA when your workflow does not benefit much from it because the state at t depends on the state at t-1. But the improvement of these features is definitely not negligible ; at the beginning of the SPECTRE / Meltdown / ... mitigations the loss of performance was double-digit big% in some cases.
- adgjlsfhk1 5y agoThis isn't really true. Micro-op caches are fairly new, branch predictors are massively improved, caches have gone from 1 level to 3, lots of operations have gotten way more efficient (64 bit division for example has gone from around 60 cycles to 30 cycles between 2012 and now). Out of order execution has also massively improved, which allows for major speed increases.
- jcelerier 5y agoL3 caches have been in consumer Intel CPUs since 2008 and uop caches were already there in pentium 4 (released in 2000, almost 22 years ago :-)). Hardly new. Of course there are interesting iterative improvements, but nothing earth-shattering.
- adgjlsfhk1 5y agoYou might note that neither 2008 nor 2000 are the 1980s which was the time you previously referred to.
- jcelerier 5y agodouble-checked and L3 was actually also there in P4 in 2003 ; and P4 itself was in the works since 1998. For me that's closer to late 90s (which is also what I referred to) than today, that's almost as many years as there were between the two world wars...
- qayxc 5y ago> But if that's the case, what's the point in all that extra-complexity in the CPU, if in the end the benefits seem to be miniscule? They aren't. 10 years ago, single thread performance was achieved by upping the core frequency. That trend died when it hit physical limitations and we're stuck with 4-5 GHz ever since. In order to get more performance, all these tricks (caches, speculative execution, data-parallelism, etc.) had to be employed in addition to more cores. In audio processing this means that a modern laptop can process more effects and tracks than a beefy workstation could in 2011. Sure, each single effect still taxes the CPU pretty bad; but in contrast to 2011 this means you can easily run dozens in parallel without breaking a sweat or endless fiddling with buffer sizes to keep latency and processing capability in balance.
- pjc50 5y ago> we need processors that can do high throughput on many execution units, and languages that support that well. Language support is very difficult for this, for a whole bunch of reasons - and it often requires redesigning the entire program and its data structures. It is still the case that most code the end-user is waiting for is JITted Javascript, which is why Apple focused so much effort into making that fast. Which is forced to be single-threaded. Hence all the big/little CPU designs; you get one or two high-speed high-power cores, and some low-speed low-power cores.
- usrbinbash 5y agogo doSomethingWith(x) Threads.@threads for x = 1:42 thread::spawn How is that difficult? The problem is that people are taught that concurrency and the capacity for parallell execution is somehow difficult. It really isn't. > It is still the case that most code the end-user is waiting for is JITted Javascript That's a problem with JavaScript, not with language design. JS is simply not a very good language, and its lack of support of parallell processing is just one of its many problems.
- saagarjha 5y agoIt's difficult for several reasons, and you've identified one of them: people aren't taught how to write code that can take advantage of concurrency. Except…this includes you. I work as a performance engineer, and a lot of my job is actually undoing concurrency written by people who do it incorrectly and create problems worse than they could've ever have without it. People will farm work out to a bunch of threads, except they'll have the work units be so small that the synchronization overhead is an order of magnitude more than the actual work being done. They'll create a thread pool to execute their work and forget to cap its size, or use an inappropriate spawning heuristic, and cause a thread explosion. They'll struggle mightily to apply concurrency to a problem that doesn't parallelize trivially, due to involved data dependencies, and write complex code with subtle bugs in it. Writing concurrent code is hard. In general, nobody actually wants concurrency*, it's just a thing we deal with because single-threaded performance has stopped advancing as fast as we'd want it to. As an industry we're slowly getting more familiar with it, and providing better primitives to harness it safely and efficiently, but the overall effort is a whole lot harder than just slapping some sort of concurrent for loop around every problem. *Except for some very rare exceptions that cannot be time shared
- eloff 5y agoThere are already processors that are highly parallel, high throughput execution. GPUs as you point out. The reason they haven't replaced CPUs is not laziness, it's that many problems are not trivially parallel. Today's mix of CPUs that are fast at serial execution plus multiple cores, SIMD, and GPUs seems to give a good mix of flexibility to program for.
- IshKebab 5y agoRust and Go don't support GPU-style parallelism natively (I guess Julia probably does). You wouldn't even want that for 99% of programs. It's only useful for big mathsy tensor operations with very little flow control. Most programs are not like that at all.
- pyjarrett 5y ago> languages that support that well Ada has had built-in tasking for a long time, and is now also getting a built-in parallel blocks and parallel loops (iterators and for) structures in Ada 2022.