18 ms·
If cache coherence is relevant to you, I strongly recommend the book “A Primer on Memory Consistency and Cache Coherence”. It’s much easier to understand the de
by strstr 7y ago
If cache coherence is relevant to you, I strongly recommend the book “A Primer on Memory Consistency and Cache Coherence”. It’s much easier to understand the details of coherency from a broader perspective, than an incremental read-a-bunch-of-blogs perspective.
I found that book very readable, and it cleared up most misconceptions I had. It also teaches a universal vocabulary for discussing coherency/consistency, which is useful for conveying the nuances of the topic.
Cache coherence is not super relevant to most programmers though. Every language provides an abstraction on top of caches, and nearly every language uses the “data race free -> sequentially consistent”. Having an understanding of data races and sequential consistency is much more important than understanding caching: the compiler/runtime has more freedom to mess with your code than your CPU (unless you are on something like the DEC Alpha, which you probably aren't).
If you are writing an OS/Hypervisor/Compiler (or any other situation where you touch asm), cache coherence is a subject you probably need a solid grasp of.
- jblow 7y agoDisagree on that last part. If more programmers understood cache coherency, maybe their programs would not run like a giant turd.
- codetrotter 7y agoI agree with you Jonathan but am wondering, will Jai help programmers write programs with better cache coherency even if said programmers don’t understand cache coherency well? Or is that orthogonal to the goals of Jai?
- cma 7y agoIf it still plans on wrapping SOA in something that looks in use like AOS, it could make people less aware of how their code is impacting cache (would now have to look at the definitions instead of seeing the array form at point of usage). But if it is enough more ergonomic to write AOS code it might still be worth it and increase uptake.
- Sean1708 7y agoThe developer would still need to understand enough to be able to choose between them though, even if using them is identical.
- naikrovek 7y agoLike any reasonable language, I imagine JAI will allow a developer to write software with the cache in mind. It will not force them to, nor will it prevent them from doing otherwise.
- BubRoss 7y agoOnly if he releases it
- strstr 7y agoMost engineers don't write code with hard performance constraints. Game devs probably need to be fighting to get every frame. For the bulk of the eng I work with the concept of StoreLoad reordering on x86 would be an academic distraction.
- kasey_junk 7y agoThe problem is you never know when you’ll go from “dev who doesn’t care” to “dev who does”. While I agree that the details of StoreLoad are likely a distraction the big picture concepts of cache coherence presented in this article are table stakes for performant systems.
- strstr 7y agoWhile that might be true in practice, I do think not knowing when you transition between those two indicates a failing by your coworkers/mentors. I encourage people I work with to read many of the books in this book series. I particularly encourage them to read “Hardware and Software Support for Virtualization”, since it’s basically a book on their job.
- steev 7y agoJust to be clear, the book series you are referring to is "Synthesis Lectures on Computer Architecture" published by Morgan-Claypool, right? If you had to rank them in importance for the average engineer, how would you rank them?
- onion2k 7y agoThe problem is you never know when you’ll go from “dev who doesn’t care” to “dev who does”. This is true for literally every hard problem in dev though, and the implication that you need to grasp everything just in case you need it is silly. The problem space in compsci is too big to know everything. We have to choose.
- afiori 7y ago
- fulafel 7y agoAre the details really helpful for performance work? MOSI, MOESI, MERSI, MESIF etc are 99% irrelevant to having the right metal model. "Dirtying cache lines across different cores/threads is slow" is most of what you need to know. Within the same core the coherency protocols between levels of cache is not really visible at all to software even as varying performance artifacts. Of course you might end up analyzing assembly level perf traces in a hot path for some game console with a less known CPU architecture and making a cross cpu cache miss slightly less slow could just maybe be helped by the detailed understanding of the machine model, but by that time you're already far in the not-giant-turd territory (at least if you're optimizing the right thing). Of course computer architecture is fascinating and fun to learn about.
- vardump 7y ago> Of course you might end up analyzing assembly level perf traces in a hot path for some game console with a less known CPU architecture Or if you work in finance, mining, oil industry, medical, biosciences and countless other fields where you need to get good performance out of your hardware. Yes, even if you use GPUs, they're no magic bullet, they also have their architectural bottlenecks, strengths and weaknesses. Or if you care about power consumption. There are a lot of reasons to optimize hot paths and inner loops. CPU single core performance isn't improving much anymore, and we need to make better use of what we have.
- fulafel 7y agoI was trying to say that even in those cases you almost never benefit from knowing the details of cache coherency.
- vardump 7y agoThat's a very bold statement. You could just as well for example say that front end Javascript developer almost never needs to understand event callbacks or how DOM works. If you write multithreaded high performance code, yeah, you do need to know about cache coherency at varying levels of detail. Sometimes rough rules of thumb work, not too often you need to understand all those annoying performance destroying details that leak through cache abstraction.
- jacobush 7y agoMaybe but it feels like most are stuck in environments which will do bad things to their cache coherence. It's fine if you are doing some data processing in C. If you are using .NET with a bunch of magic libraries or Javascript or whatever, sure it will help, but to actually make an impact you have to be very careful.
- devnonymous 7y agoI disagree on this. More programmers should understand the performance characteristics of the abstraction layers that they rely on. Else we have the case of jerk programmers who insist on redesigning / rewriting / refactoring to optimise for CPU caches while still running the apps via a bunch of docker containers each based of the centos image to run one simple binary that probably needs only glibc.
- DrScientist 7y agoSurely the whole point of good system design is a set of logical abstractions, where you need to understand the logical model and not the internal details - as these are free to be evolved. Of course performance matters, but surely having performance tests, rather than trying to second guess what the whole stack below you might be doing, is 1. more efficient 2. more accurate 3. more likely to detect changes in a timely way. That's not to say, you shouldn't be curious and deep understanding isn't a good thing. Just saying understanding inside-out the abstraction you are working with ( eg Java Memory Model ) it's performance characteristics ( from real world testing ) - is more important than some passing knowledge of real world CPU design. This app I am using right now is in a webbrowser - not sure how understanding cache coherency helps in a single threaded javascript.
- skohan 7y ago> Surely the whole point of good system design is a set of logical abstractions, where you need to understand the logical model and not the internal details In my opinion an important part of being a good programmer is understanding - at least at a broad level - how the set of abstractions you're working on top of work. At the end of the day, our job is making computer hardware operate on data in memory. The more that we forget that, and think about computing as some abstract endeavor performed in Plato's heaven, the more tendency we have for bloat and inefficiency to creep into our various abstractions. In other words, I think it's better to think about abstractions as a tool for interacting with hardware, not as something to save us from dealing with hardware.
- pingyong 7y agoEh, idk. Most programs that run like a giant turd do it because they load 15 megabytes of Javascript libraries to call two functions, or something to that effect. Computers are so fast now that you really need to be doing something unbelievably stupid for things in consumer programs to not be instantaneous.
- claudius 7y agoLike start Firefox? I have no idea why it takes multiple seconds to bring up a blank window when starting a comparably useful program like Claws Mail is absolutely instantaneous.
- pingyong 7y agoTo be honest, yes, generally programs that don't start instantaneously do something very unnecessary and stupid that has nothing to do with actual CPU performance. As in they read thousands of tiny files or they decompress files or they wait for some sort of answer from the network, or they make 10k+ expensive API calls or all of the above. Firefox is so big it might fall into the "all of the above" category, but to answer that question definitively you'd have to analyze what Firefox actually does. And of course, making the decompression 15-20% faster by optimizing the decompression code (which is usually not even written by the developers of said software but just some external library) won't even make a difference because 20% less than 5 seconds is still 4 seconds which is way too long for a program to start. Instead using a different compression algorithm that increases file size by 25% but decompression speed by 10x would actually start solving the problem, with the next step being to ask why the program needs to read so much damn data at the start in the first place. But since NVMe SSDs and Intel CPUs with very high boost clocks are quickly becoming the norm now even for laptops I don't see much of that happening, because Firefox starts pretty quickly (~1 second) on those machines.
- bzbarsky 7y agoFwiw, Firefox stores almost everything it will need on startup in one large file (omni.ja) precisely to avoid the "thousands of tiny files" problem. That data is uncompressed, precisely to avoid the decompression problem. The low-hanging fruit has largely been picked. As for why so much data needs to be read... I just checked, and on Mac the main Firefox library (the executable itself is mostly a stub) is 120MB. So that's going to take a second or three just to read in at typical HDD speeds (faster on a good SSD), and then the dynamic linker has to do its thing on that big library, which is not instantaneous either.
- voldacar 7y agoCould you explain what the DEC Alpha did differently here? It was before my time :)
- ridiculous_fish 7y agomemory-barriers.txt from Linux is a lovely way to be introduced to memory barriers in general, and the Alpha memory model in particular. https://github.com/torvalds/linux/blob/master/Documentation/memory-barriers.txt https://github.com/torvalds/linux/blob/master/Documentation/...
- rawoke083600 7y agoMan I just LOVE how the linux kernel can have such detailed and important information in a simple TEXT file. Fantastically functional and prove they focusing on the right stuff. At my previous company. I would have to spend the better part of a day to get my "documentation" in the "correct" Confluence Style And Manner. They were adamant THAT'S were the value is, to have documentation is the most beautiful and absurd style" and double-linked structure. You would have to block out a day or two in your scrum(what nonsense) just to focus on your documentation. And this is not some sort of important* software like Linux or Banking... but a stupid website.
- brandmeyer 7y agoIMO, the various blogs and tutorials out there that help to make sense of the C++11 memory model make better tutorials than the Linux kernel's own shenanigans. The C++11/C11 memory model added memory_order_consume specifically to support the Alpha. https://preshing.com/20140709/the-purpose-of-memory_order_consume-in-cpp11/ https://preshing.com/20140709/the-purpose-of-memory_order_co...
- gpderetta 7y ago> The C++11/C11 memory model added memory_order_consume specifically to support the Alpha. yes and no. Alpha is relevant because it is the only architecture where consume requires an explicit barrier, but then again, I think the revised C++11 memory model might not even be fully implementable on Alpha; Consume primarily exist because acquire is very expensive to implement in traditional RISCs like POWER and 32 bit ARM, while consume can be implemented with a data or control dependency. Aarch64 has efficient acquire/release barriers, so it is less of an issue. /pedantic
- deepaksurti 7y ago“A Primer on Memory Consistency and Cache Coherence” is part of the "Synthesis Lectures on Computer Architecture" which are 50-100 page booklets on topics related to HW components. All the booklet PDF's are available online [1]. edit: only those PDF's with a checkmark are available as PDF to download, the rest can be bought. Quite a few actually available for download. [1] https://www.morganclaypool.com/toc/cac/1/1 https://www.morganclaypool.com/toc/cac/1/1
- kqr 7y agoYou make it sound like CPU caches are the only caches around. I deal with higher-level caching a lot, and I'm not writing an OS. Is your book recommendation still useful for me?