5 ms·
Yeah, given a predictable access pattern, latency can be completely hidden by prefetching. For an unpredictable access pattern, there's a physically imposed lo
by sparky 13y ago
Yeah, given a predictable access pattern, latency can be completely hidden by prefetching. For an unpredictable access pattern, there's a physically imposed lower bound on random access latency, roughly proportional to the square root of the number of bits stored.
There's a huge tradeoff space between latency, bandwidth, capacity, and cost, and market forces have forced a convergence around 2 design points, colloquially DDR and GDDR. For yield reasons, die area (cost) has been mostly fixed for a long time. Successive generations of DDR spend the dividends of Moore's Law primarily on extra capacity, then on a bit of additional bandwidth when possible. Successive of generations of GDDR prioritize bandwidth, primarily by dedicating tons of area to high-speed single-ended I/Os.
These two design points make sense for their most common use cases. In traditional disk-based systems, avoiding hitting the disk is more important than absolute DRAM latency, so increasing capacity is your best bet. On GPUs, you need enough bandwidth to feed a quickly growing number of functional units on the chip, and at least for graphics, the access pattern can be made to be extremely predictable, so latency is not as important there.
The advent of faster-than-HDD persistent storage (SSDs) and the desire to run more general purpose workloads on highly parallel machines like GPUs points to a need for a third DRAM design point.
- wmf 13y agoThere is RLDRAM, but I've heard that it's expensive.
- sparky 13y agoYeah, there are a few specialized parts for networking, telecom equipment, defense, and other less cost-sensitive applications, but DDR and, to a lesser extent, GDDR dominate the market and enjoy much greater economies of scale.
- marshray 13y agoPerhaps another factor is the software's difficulty to successfully utilize more than a couple of cores? I imagine Intel thinking "We're a couple of process steps ahead of everyone else, we should take advantage of that. But adding more cores in the same package has reached dimishing returns due to pin count and main memory bandwidth. Most applications won't utilize that 13th hyperthread anyway. So let's improve the memory subsystem from within the package."