5 ms·
> Note that we’re 40 years of Moore’s law scaling later and the available memory bandwidth per instruction has gone down substantially. This is unfair, and (th
by pslam 9y ago
> Note that we’re 40 years of Moore’s law scaling later and the available memory bandwidth per instruction has gone down substantially.
This is unfair, and (this is me being unfair now) this article is missing the woods for the trees.
There were engineering pressures which resulted in the current ratios. I think it is fairer to say that the current situation — where there's about (hand-waving) 1 byte per instruction of bandwidth per core — reflects the kinds of tasks we expect our machines to be doing. It is very rare to find a task which is memory speed bound. There's almost always substantial processing to be done with data.
It's not even that hard to increase memory bandwidth. You "just" double up memory channels. This is of course expensive, which in turn is a back-pressure which results in architectures designed around the current sweet-spot.
I'm also puzzled the author thinks the situation is "worse". Pretty much every desktop class machine I've used from about 1990-2005 was extremely starved of memory bandwidth, and cores did a far worse job of hiding latency (OOO renames etc). What we have today feels fairly comfortable, to me at least, with some outlier tasks where you might want more (and then obtain specialist hardware).
This is a long-winded way of saying: the current core vs memory speed ratio is a sweet-spot of cost vs efficiency, and works well given the tasks and algorithms we execute on these machines. What we had in the olden days was just a case of unoptimized architecture, which hadn't converged yet.
- vvanders 9y ago> It is very rare to find a task which is memory speed bound. There's almost always substantial processing to be done with data. One could argue that memory speed doesn't matter because memory latency has remained (relatively) constant since the advent of DDR. Can't process something while you're waiting for that cache miss to complete.
- nialo 9y agoI don't think it's worse. The (possibly missing) context of this post is a Compute Shader running on a GPU spending more time writing out the results of a computation than actually doing the computation. As I read this post, the moral of the story is just to try to write your code in such a way as to always have that substantial work to do with the data. Prefer one big pass over everything to several smaller passes, each of which must write out results, especially on a GPU. Try to actually have 11 instructions per byte of memory access. I don't think the intent is to make any argument in particular about the state of CPU or GPU design or anything of the sort.