12 ms·
I see nobody pointing out the relatively low number of integer registers (32 64-bit registers [1]) for a VLIW that is intended to emulate an high-end Out-of-Ord
by avianes 4y ago
I see nobody pointing out the relatively low number of integer registers (32 64-bit registers [1]) for a VLIW that is intended to emulate an high-end Out-of-Order machine.
Modern high-end Out-of-Order processors feature over 160 physical integer registers.
This number is high because OoO processors use a very large instruction window, and every instruction in this window have to be free from register name false-dependency (Write-after-Read and Write-after-Write dependency). That's the register-renaming job to introduce extra physical register to hide false-dependency.
If they really emulate the behavior of an OoO processor we should expect more registers.
The large instruction window form OoO high-end processor provides latency tolerance, with VLIW processors, to increase latency tolerance, we usually increase the register count and employ more aggressive static scheduling techniques.
It is known that register-renaming technique allocates on average more registers than necessary, due to early allocation and late release, so with static scheduling it should be possible to make more efficient allocate/release.
But here the register count difference is too large, low register allocation/release efficiency of register-renaming is not enough to explain the difference.
As a comparison, Transmeta VLIW core has 64 integer register and 48 integer shadow-register which is equivalent to the physical register count of OoO processors of its time.
So either they don't really emulate an out of order processor, or they do some magic, or they don't really perform that well..
(or documents that mention 32 integer registers are incorrect)
The most likely hypothesis to me is that their benchmarks numbers are based on a selection of benchmarks that can exploit their vector unit, and other uncompetitive bench results have been left out.
Meaning that performance on typical x86 general purpose applications would be extremely low with this processor.
[1] https://www.tachyum.com/media/pdf/tachyum-tries-for-hyperscale-servers-2.pdf https://www.tachyum.com/media/pdf/tachyum-tries-for-hypersca...