4 ms·
This is a good set of slides. Dan is a good guy. There are a few nits I would pick. Sqrt(N) convergence comes from independence not normality -- based on ind
by cb321 1y ago
This is a good set of slides. Dan is a good guy. There are a few nits I would pick. Sqrt(N) convergence comes from independence not normality -- based on independence => linearity of variance. { So, N IID samples of any distribution have a sum with N times higher variance, but then dividing by N you et sqrt(N). } There is, of course, a higher order relationship between the variance / "scale^2" of the distro and its tails which statisticians refer to as "shape". He later goes on to mention the dependence problem, though, and the minimum dt solution that, relied upon by, e.g., https://github.com/c-blake/bu/blob/main/doc/tim.md https://github.com/c-blake/bu/blob/main/doc/tim.md. So, it's all good. He may have covered it in voice, even.
He also mentions the Sattolo used by https://github.com/c-blake/bu/blob/main/doc/memlat.md https://github.com/c-blake/bu/blob/main/doc/memlat.md to do his memory latency measurements. One weird thing was how he said because of 1 byte/cycle is 4GB/s things are "easily CPU bound" while I feel like I've been "fighting The Memory Wall for at least 3 decades now..." even just from super-scalar CPUs, but he later does some vectorization stuff. That more relates to what calcs you are doing, of course, but high bandwidth memory is a big part of what nVidia is selling.
- mananaysiempre 1y ago> One weird thing was how he said because of 1 byte/cycle is 4GB/s things are "easily CPU bound" while I feel like I've been "fighting The Memory Wall for at least 3 decades now..." even just from super-scalar CPUs I mean, he’s right? For mostly-linear processing like Lemire’s projects that he mentions in his talk (simdutf, simdjson), memory latency is completely absorbed by the prefetcher, so only memory bandwidth matters. You can’t saturate DDR4 bandwidth, or even a tenth of DDR5 bandwidth (or, like Lemire is fond of mentioning, even PCIe Gen 4 SSD bandwidth) if your code (after autovectorization) is structured around doing a thing per byte. Probably not even a thing per UTF-16 code unit. So in those cases—no, you’re not hitting the memory wall. Not even close. You’re hitting the clock-speed wall.
- cb321 1y agoWell, like I said and you said - it depends on the calculations. :-) FWIW, I never meant to suggest he was wrong in any absolute sense, but I did do a double take reading that. Besides the fine distinction you made about byte-at-a-time, one other way I sometimes try to express the variety is "BLAS L1/L2/L3". These levels roughly correspond to how much "loop nesting" there is, or roughly how well can we amortize the cost of memory transfers. So, L1 would be scaling each element by 10.0X, say, while L3 would be matrix multiply. One might have to be steeped in numerical linear algebra to appreciate that kind of analogy/terminology, though. Latency can also be a big issue vs. BW in "it all depends." Ultimately, it comes down to "CPUs can do A LOT per clock cycle" these days - multi-way issue of up to 64-way vector instructions (byte-wise on avx512) and so on and what 1 cycle means can just vary by orders of magnitude (not even including multi-core orders of magnitude and then distributed orders of magnitude). But "not always". And it all depends. Part of CPUs getting so "big" is that the "play/drift" in various statements has a lot more flexibility. So, cross-talk about issues like this has also increased. As have, probably, double takes like I mentioned. :-)