6 ms·
If main memory is so slow, would it be better to treat cache as memory and handle it directly in code rather than try to trick it into doing what we want?
by nulltype 10y ago
If main memory is so slow, would it be better to treat cache as memory and handle it directly in code rather than try to trick it into doing what we want?
- chris_va 10y ago__builtin_prefetch sort of works in x86, but you'll probably end up doing more harm than good when second guessing the automated cache.
- nulltype 10y agoRight, I'd like to avoid that complexity entirely, just get direct access to things rather than hinting what I want.
- newobj 10y agoThe cache is a shared resource, so I'm not sure how you would propose to "handle it directly", other than literally having a MMU/virtual memory layer on top of it, and try to provide some guarantees about coherence. But given how small it is, without allowing for hard pins, I suspect that the entire cache would be thrashed pretty badly across processes and it will just degrade back to what happens today; it's useful/coherent for short/co-located bursts and not much else.
- nulltype 10y agoWell imagine I could address memory on the CPU die like a register but say 1MB of memory, located right next to the normal place you put a cache. It does not have to be coherent I think. Process thrashing sounds like an issue though. I imagine if the address space is across all CPU cores then the software can pin to a core so that the access is local. Remote access to another CPU will be slow, but I'm not sure how great it is with caches anyway.
- corysama 10y agohttps://en.wikipedia.org/wiki/Scratchpad_memory https://en.wikipedia.org/wiki/Scratchpad_memory On the original PlayStation, the scratchpad was a PITA. You had to weigh the speedup against the cost of manually copying in and out the the pad. A modern implementation would copy via DMA like the PS3's Cell processors. Still a PITA, but at least the payoff for the effort can be quite good.
- nulltype 10y agoNice! This is exactly what I was thinking of, thanks for the link.
- clevernickname 10y agoCame here to post this. The other advantage of cache over explicit control of the on-chip SRAM is that it allows code to be inherently forwards and backwards compatible, rather than tying the code to a specific machine with a specific amount of on-chip memory. And I imagine the nightmare would only increase if you had to factor multitasking into an explicit scheme. The primary benefit of explicit on-chip memory is not, AFAIK, that you can manually manage it significantly better than a cache, but that it takes up significantly less die space and has lower access latency. You really see this idea shine in tiny cheap microcontrollers that have no external memory.