5 ms·
Not sure if it will be available, but 400GB/s is way too much for 8 cores to take up. You would need some sort of avx512 to hog up that much bandwidth. Moreove
by znwu 5y ago
Not sure if it will be available, but 400GB/s is way too much for 8 cores to take up. You would need some sort of avx512 to hog up that much bandwidth.
Moreover, it's not clear how much of a bandwidth/width does M1 max CPU interconnect/bus provide.
--------
Edit: Add common sense about HPC workloads.
There is a fundamental idea called memory-access-to-computation ratio. We can't assume a 1:0 ratio since it was doing literally nothing except copying.
Typically your program needs serious fixing if it can't achieve 1:4. (This figure comes from a CUDA course. But I think it should be similar for SIMD)
Edit: also a lot of that bandwidth is fed through cache. Locality will eliminate some orders of magnitudes of memory access, depending on the code.
- GeekyBear 5y agoA single big core in the M1 could pretty much saturate the memory bandwidth available. https://www.anandtech.com/show/16252/mac-mini-apple-m1-tested https://www.anandtech.com/show/16252/mac-mini-apple-m1-teste...
- lowbloodsugar 5y agoDon't know the clock speed but 8 cores at 3Ghz working on 128bit SIMD is 8316 = 384GB/s so we are in the right ball park. Not that I personally have a use for that =) Oh, wait, bloody Java GC might be a use for that. (LOL, FML or both).
- dragontamer 5y agoBut the classic SIMD problem is matrix-multiplication, which doesn't need full memory bandwidth (because a lot of the calculations are happening inside of cache). The question is: what kind of problems are people needing that want 400GB/s bandwidth on a CPU? Well, probably none frankly. The bandwidth is for the iGPU really. The CPU just "might as well" have it, since its a system-on-a-chip. CPUs usually don't care too much about main-memory bandwidth, because its like 50ns+ away latency (or ~200 clock ticks). So to get a CPU going in any typical capacity, you'll basically want to operate out of L1 / L2 cache. > Oh, wait, bloody Java GC might be a use for that. (LOL, FML or both). For example, I know you meant the GC as a joke. But if you think of it, a GC is mostly following pointer->next kind of operations, which means its mostly latency bound, not bandwidth bound. It doesn't matter that you can read 400GB/s, your CPU is going to read an 8-byte pointer, wait 50-nanoseconds for the RAM to respond, get the new value, and then read a new 8-byte pointer. Unless you can fix memory latency (and hint, no one seems to be able to do so), you'll be only able to hit 160MB/s or so, no matter how high your theoretical bandwidth is, you get latency locked at a much lower value.
- lillecarl 5y agoDoesn't prefetching data into the cache more quickly assist in execution speed here?
- dragontamer 5y agoHow do you prefetch "node->next" where "node" is in a linked list? Answer: you literally can't. And that's why this kind of coding style will forever be latency bound. EDIT: Prefetching works when the address can be predicted ahead of time. For example, when your CPU-core is reading "array", then "array+8", then "array+16", you can be pretty damn sure the next thing it wants to read is "array+24", so you prefetch that. There's no need to wait for the CPU to actually issue the command for "array+24", you fetch it even before the code executes. Now if you have "0x8009230", which points to "0x81105534", which points to "0x92FB220", good luck prefetching that sequence. -------- Which is why servers use SMT / hyperthreading, so that the core can "switch" to another thread while waiting those 50-nanoseconds / 200-cycles or so.
- lillecarl 5y agoI don't really know how the implementation of a tracing GC works but I was thinking they could do some smart memory ordering to land in the same cache-line as often as possible. Thanks for the clarifications :)
- monocasa 5y agoInterestingly earlyish smalltalk VMs used to keep the object headers in a separate contiguous table. Part of the problem though, is that the object graph walk pretty quickly is non contiguous, regardless of how it's laid out in memory.
- kaba0 5y agoBut that’s just the marking phase, isn’t it? And most of it can be done fully in parallel, so while not all CPU cores can be maxed out with that, more often than not the original problem itself can be hard to parallelize to that level, so “wasting” a single core may very well be worth it.
- deleted 5y ago[deleted]
- terafo 5y ago> Not sure if it will be available, but 400GB/s is way too much for 8 cores to take up. You would need some sort of avx512 to hog up that much bandwidth. If we assume that frequency is 3.2Ghz and IPC of 3 with well optimized code(which is conservative for performance cores since they are extremely wide) and count only performance cores we get 5 bytes for instruction. M1 supports 128-bit Arm Neon, so peak bandwidth usage per instruction(if I didn't miss anything) is 32 bytes.
- deleted 5y ago[deleted]