3 ms·
I'm not entirely sure about the backend, but for the frontend Intel has come up with a 'split decoder'. Essentially, Skymont will have 3x 3-wide decoders workin
by ColonelPhantom 2y ago
I'm not entirely sure about the backend, but for the frontend Intel has come up with a 'split decoder'. Essentially, Skymont will have 3x 3-wide decoders working in parallel on different parts of the instruction stream. I have no idea how they find points to perform load-balancing between these clusters (outside of taken branches), but the 2x3-wide setup in Gracemont seems to work like a charm after some improvements over Tremont which debuted the split decoder scheme.
> This scheme appeared with Tremont, where it could only switch decoders at taken branch boundarie and Gracemont improves load balancing between the two decode clusters by automatically switching, instead of relying on taken branches. That helps performance in very long unrolled loops, which could get stuck on one decode cluster on Tremont. When we tested with longer loop lengths, we didn’t see any drop off in Gracemont’s instruction throughput:
- https://chipsandcheese.com/2021/12/21/gracemont-revenge-of-the-atom-cores/ https://chipsandcheese.com/2021/12/21/gracemont-revenge-of-t...