3 ms·
> DRAM has 50+ ns of latency no matter how close you put it to the chip. This is not some fundamental law; it's entirely dependent on bitline and wordline capa
by sparky 13y ago
> DRAM has 50+ ns of latency no matter how close you put it to the chip.
This is not some fundamental law; it's entirely dependent on bitline and wordline capacitance (array size). For external DRAM chips, huge arrays make sense, and you end up with tens of ns of latency. In an eDRAM cache scenario, you use much smaller arrays, and in the case of the POWER7 L3, latency is more like 6ns (for row hits, of course).
- marshray 13y agoIt seems to me (you may know more about this than I do) that one can design a memory system to make latency for sequential addresses accesses arbitrarily small and this is mostly how (G)DDR(n) have improved overall performance. But random (row miss) accesses don't seem to have improved much. From early 1980's to early 2000's, DRAM latency went from 150 to 50 ns and has stayed around there since. That's 1-2 doublings of performance in that parameter in the last 30 years, compared to who-knows how many in transistor size and speed. Still, Intel knows what they're doing and I can see a place for (yet another) spot in the memory hierarchy before the CPU resort to off-package DRAM.
- sparky 13y agoYeah, given a predictable access pattern, latency can be completely hidden by prefetching. For an unpredictable access pattern, there's a physically imposed lower bound on random access latency, roughly proportional to the square root of the number of bits stored. There's a huge tradeoff space between latency, bandwidth, capacity, and cost, and market forces have forced a convergence around 2 design points, colloquially DDR and GDDR. For yield reasons, die area (cost) has been mostly fixed for a long time. Successive generations of DDR spend the dividends of Moore's Law primarily on extra capacity, then on a bit of additional bandwidth when possible. Successive of generations of GDDR prioritize bandwidth, primarily by dedicating tons of area to high-speed single-ended I/Os. These two design points make sense for their most common use cases. In traditional disk-based systems, avoiding hitting the disk is more important than absolute DRAM latency, so increasing capacity is your best bet. On GPUs, you need enough bandwidth to feed a quickly growing number of functional units on the chip, and at least for graphics, the access pattern can be made to be extremely predictable, so latency is not as important there. The advent of faster-than-HDD persistent storage (SSDs) and the desire to run more general purpose workloads on highly parallel machines like GPUs points to a need for a third DRAM design point.
- wmf 13y agoThere is RLDRAM, but I've heard that it's expensive.
- sparky 13y agoYeah, there are a few specialized parts for networking, telecom equipment, defense, and other less cost-sensitive applications, but DDR and, to a lesser extent, GDDR dominate the market and enjoy much greater economies of scale.
- marshray 13y agoPerhaps another factor is the software's difficulty to successfully utilize more than a couple of cores? I imagine Intel thinking "We're a couple of process steps ahead of everyone else, we should take advantage of that. But adding more cores in the same package has reached dimishing returns due to pin count and main memory bandwidth. Most applications won't utilize that 13th hyperthread anyway. So let's improve the memory subsystem from within the package."