3 ms·
Two questions: 1) How does a core know in which private cache of another core it can put the evicted data? Wouldn't that just remove data from that other core'
by hesk 5y ago
Two questions:
1) How does a core know in which private cache of another core it can put the evicted data? Wouldn't that just remove data from that other core's cache that the other core needs?
2) Cores in today's processors can already read data from the private cache of another core (i.e., snooping), can they not?
- whack 5y ago> How does a core know in which private cache of another core it can put the evicted data? Good question. There are many ways of handling this and I'm curious how IBM decided to implement this. Though it's unlikely they will divulge such internal implementation details. > Wouldn't that just remove data from that other core's cache that the other core needs? "This becomes important for cloud services (yes, IBM offers IBM Z in its cloud) where tenants do not need a full CPU, or for workloads that don’t scale exactly across cores." > Cores in today's processors can already read data from the private cache of another core (i.e., snooping), can they not? Depends on the exact CPU and its implementation. But yes, this idea of a "private cache" is pretty silly and misleading. Each cache is just a small subset of the system's memory. And all cores can issue memory reads which will eventually get the data from somewhere. Whether you're reading it from your own private cache, or some other core's private cache, a shared cache, or from DRAM, does not change the final result. The only distinction here is in performance. Accessing a cache that is colocated to your core is much faster. Accessing a cache that is either shared or located elsewhere, is much slower. An approximate analogy would be using a us-east-1 ec2 instance to read from a us-east-1 RDS database, as opposed to a us-west-2 RDS database. The former database and its contents are certainly not "private", but they are a whole lot faster.
- ethbr0 5y agoI found the easiest way to think about it, when we were doing caches in computer architecture was to think of it in terms of abstraction. The CPUs have "read memory" and "write memory" as the only operations programs can see. The cache policies operate asynchronously (because multiple cores simultaneously executing), to make those happen in the fastest way possible. There's an endless variety of schemes and ways to design it, with the one huge requirement that it must ALWAYS be accurate. But as long as you can satisfy that, you can do some goofy, crazy stuff. Might be faster, might be slower, but it will work.
- Taniwha 5y agoThere's likely some 'recently used' state, plus they are probably only pushing dirty lines into other caches to save the overhead of pushing them all the way to memory on some miss/allocates