4 ms·
I read all that marketing stuff and my brain just sees APU. I guess at some level, that’s just marketing stuff too, but it’s not a new idea.
by binary132 2y ago
I read all that marketing stuff and my brain just sees APU. I guess at some level, that’s just marketing stuff too, but it’s not a new idea.
- sroussey 2y agoEh… not quite. Maybe on an Instinct. Unified memory means the CPU and CPU means they can do zero copy to use the same memory buffer. Many integrated graphics segregate the memory into CPU owned and GPU owned, so that even if data is on the same DIMM, a copy still needs to be performed for one side to use what the other side already has. This means that the drivers, etc, all have to understand the unified memory model, etc. it’s not just hardware sharing DIMMs.
- binary132 2y agoI was under the impression PS4’s APU implemented unified memory, and it was even referred to by that name[1]. APUs with shared everything are not a new concept, they are actually older than programmable graphics coprocessors… https://www.heise.de/news/Gamescom-Playstation-4-bietet-Unified-Memory-Xbox-One-nicht-1939716.html https://www.heise.de/news/Gamescom-Playstation-4-bietet-Unif...
- sunshowers 2y agoI believe that at least on Linux you get zero-copy these days. https://www.phoronix.com/news/AMD-AOMP-19.0-2-Compiler https://www.phoronix.com/news/AMD-AOMP-19.0-2-Compiler
- sliken 2y agoThe new idea is having 512 bit wide memory instead of PC limitation of 128 bit wide. Normal CPU cores running normal codes are not particularly bandwidth limited. However APUs/iGPUs are severely bandwidth limited, thus the huge number of slow iGPUs that are fine for browsing but terrible for anything more intensive. So apple manages decent GPU performance, a tiny package, and great battery life. It's much harder on the PC side because every laptop/desktop chip from Intel and AMD use a 128 bit memory bus. You have to take a huge step up in price, power, and size with something like a thread ripper, xeon, or epyc to get more than 128 bit wide memory, none of which are available in a laptop or mac mini size SFF.
- reliabilityguy 2y ago> instead of PC limitation of 128 bit wide Memory interface width of modern CPUs is 64-bit (DDR4) and 32+32 (DDR5). No CPU uses 128b memory bus as it results in overfetch of data, i.e., 128B per access, or two cache lines. AFAIK Apple uses 128B cache lines, so they can do much better design and customization of memory subsystem as they do not have to use DIMMs -- they simply solder DRAM to the motherboard, hence memory interface is whatever they want.
- sliken 2y ago> Memory interface width of modern CPUs is 64-bit (DDR4) and 32+32 (DDR5). Sure, per channel. PCs have 2x64 bit or 4x32 bit memory channels. Not sure I get your point, yes PCs have 64 bit cache lines and apple uses 128. I wouldn't expect any noticeable difference because of this. Generally cache miss is sent to a single memory channel and result in a wait of 50-100ns, then you get 4 or 8 bytes per cycle at whatever memory clock speed you have. So apple gets twice the bytes per cache line miss, but the value of those extra bytes is low in most cases. Other bigger differences is that apple has a larger page size (16KB vs 4KB) and arm supports a looser memory model, which makes it easier to reach a large fraction of peak memory bandwidth. However, I don't see any relationship between Apple and PCs as far as DIMMS. Both Apple and PCs can (and do) solder dram chips directly to the motherboard, normally on thin/light laptops. The big difference between Apple and PC is that apple supports 128, 256, and 512 bit wide memory on laptops and 1024 bit on the studio (a bit bigger than most SFFs). To get more than 128 bits with a PC that means no laptops, no SFFs, generally large workstations with Xeon, Threadrippers, or Epyc with substantial airflow and power requirements
- Rohansi 2y agoFYI cache lines are 64 bytes, not bits. So Apple is using 128 bytes. Also important to consider that the RTX 4090 has a relatively tiny 384-bit memory bus. Smaller than the M1 Max's 512-bit bus. But the RTX 4090 has 1 TB/s bandwidth and significantly more compute power available to make use of that bandwidth.