3 ms·
I don't get it -- most of the transistors I administer are off at any given moment, way more than 50% of them. If there's heat-density issues, can't they just p
by bcoates 13y ago
I don't get it -- most of the transistors I administer are off at any given moment, way more than 50% of them. If there's heat-density issues, can't they just pack some nice relatively-cool DRAM and flash around each core and improve my bus bandwidth/latency situation?
- knappador 13y agoController latency. Hierarchy of all memory and cache has a complexity/memory-size correlation in the controller that directly affects performance. Also, power density in hot-spots is the issue, and fundamentally you want all the fastest switching parts close to each other, which fights any effort to move them apart. There will always be fast-switching, constantly used transistors next to others. Also, at the small scale, the heat has to move linearly to the chip package before it can spread three-dimensionally at all, and it's basically linear at the center of the hot-spot. The heat is going through a straw instead of a block. Multicore in essence is breaking up the hot spot and spreading it out. Our choices are limited simpler CPU's and more of them to flex this technique. GPU design is more geared toward this, and unified address space on newer AMD chips as well as potentially Nvidia's Tegra design evolution (project Denver?) both point to smaller, simpler CPU's or something like sub-CPU's (instructions are already translated to micro-opcodes and scheduled differently than they appear in the program text) that only do parts of the work instead of operating on whole threads. CPU's are binary code runtimes implemented in hardware, so this kind of abstraction is like changing the runtime without changing the bytecode fed to it. We might end up at a situation where CPU's work on 100 threads with a blurry definition of what a core even is anymore. It will happen slowly, as microprocessors retain major similarities over the years. Convergent evolution and too much engineering and experience to start anything from scratch. Can you find the FPU's in each generation? http://chip-architect.com/news/AMD_family_pic.jpg http://chip-architect.com/news/AMD_family_pic.jpg I always marveled at how much silicon is necessary for SIMD FP. I somewhat doubt this is relevant at the chip scale, but we're at 10cm per clock cycle at speed-of-light to put the frequency into perspective.
- DSingularity 13y agoTraditional DRAM process and cpu process are very different. Building eDRAM into your chip is very expensive, and raises costs significantly.