5 ms·
Totally with you on 'benchmark everything, who knows when the branch predictor will mispredict or a full cache flush will happen. (Interestingly enough, I rarel
by iheartmemcache 10y ago
Totally with you on 'benchmark everything, who knows when the branch predictor will mispredict or a full cache flush will happen. (Interestingly enough, I rarely try to beat the compiler since ICC is super smart but strategically placed, manually inserted CLFLUSHes have sped up some my work. Go figure.)
Eh L1 is bigger than you think. What cache-trashes is context switching. A common technique is shielding processes[1], delegating all interrupts to certain processors. Here's an easy example - since HT has it's own set of caches for both "CPUs" inside the physical CPU [i.e., the ALU (and all the aux AVR registers too) is shared between both], if you're doing anything computationally intensive with an unpredictable set of interrupts, but your crunch pattern is predictable enough do the following-- set the affinity for the 'cruncher' to one CPU with a full shield, and then all other components like interrupts, IO, things that you don't mind re: cache misses, etc.
[1] https://rt.wiki.kernel.org/index.php/Cpuset_Management_Utility/tutorial https://rt.wiki.kernel.org/index.php/Cpuset_Management_Utili... This is just one way to do it, but it explains the concept.