3 ms·
A load pair takes one address generation and one slot in the load queue but returns 2 x 64-bit. It doubles your load throughput at the cost of the complication
by FullyFunctional 6y ago
A load pair takes one address generation and one slot in the load queue but returns 2 x 64-bit. It doubles your load throughput at the cost of the complications from an instructions that return two results.
Apple's A14 can issues three loads per cycle. Assuming they can all be load pairs (I don't know currently) that would be 6 loads per cycle. This is frighting :)
The story is similar for store pairs, but probably less important as stores are rarely on the critical path.
- YorkshireSeason 6y agoThanks. I imagine that the complexity coming from implementing load/store pairs will probably be dwarfed by RISCV's forthcoming SIMD and Vector extensions.
- FullyFunctional 6y agoHmm, it's hard to say (and a lot depends on how performant you make the vector loads) but it's really not comparable as the load pairs target any arbitrary (and renamed) architectural registers, whereas the vector load targets a different register file and just a single vector (= much more regular).