4 ms·
I am curious about how does it compare to streaming loads? It is a common technic to avoid cache misses when you have a large number of registers, you can issue
by benou 11y ago
I am curious about how does it compare to streaming loads? It is a common technic to avoid cache misses when you have a large number of registers, you can issue loads in advance, by-passing the cache, and the CPU will manage the data dependency in HW. The idea is that hopefully when the you will use the register, its data will already be available. Otherwise you stall as usual.
- willvarfar 11y agoThere are many facets to this. From a great height they are roughly simular, but in the details they differ. For example, the Mill is a belt machine and each stack frame has its own belt and the hardware takes care of spilling in-flights across calls to any depth. Also, the Mill loads are orthogonal to cache management and we offer immunity to aliasing and false sharing.