3 ms·
You need to double the hardware... which fair enough modern CPUs are super scalar so they already do that. The issue is that there are a lot of branches in the
by TrainedMonkey 3y ago
You need to double the hardware... which fair enough modern CPUs are super scalar so they already do that. The issue is that there are a lot of branches in the code, the CPU would need to keep a lot of redundant hardware. All that extra hardware comes with increased power usage and consequently heat dissipation. On top of that modern branch predictors are pretty amazing, so you would need to get a lot of benefit to make this worth it...
So the trade is, you can get slightly better latency due to misprediction masking by executing both branches at the cost of massively decreased throughput (b.c. you are using extra hardware to execute branches of the same thread vs different threads), increased power and heat dissipation, and increased cost due to additional hardware. Note that cost, power, and usage are generally the constraints you want to satisfy, so you will generally get a significantly slower CPU that probably has worse latency due to lower clocks, less cache, and whatever other tradeoffs you need to make to fit into the power/heat/cost budget.