4 ms·
What I think is worth knowing is that compute units in GPU’s also use SMT, usually at a level of 7 to 10 threads per CU. This helps to hide latency.
by superjan 2y ago
What I think is worth knowing is that compute units in GPU’s also use SMT, usually at a level of 7 to 10 threads per CU. This helps to hide latency.
- adrian_b 2y agoMost GPUs do not use SMT, but the predecessor of SMT, fine-grained multi-threading. In every clock cycle the instruction that is initiated is selected from the many threads that are available, depending on which of them need resources that are not busy. Most GPUs do not initiate multiple instructions per clock cycle (even if multiple instructions may proceed concurrently after initiation) or if they initiate multiple instructions per clock cycle those may have to belong to distinct classes of instructions, which use distinct execution resources, for example scalar instructions and vector instructions. SMT, i.e. simultaneous multi-threading, means that in every clock cycle many instructions are simultaneously initiated from all threads and then those instructions compete for the multiple execution units of a superscalar CPU, in order to keep busy as many of those execution units as possible. For each of the concurrent execution units, e.g. for each of the 6 integer adders available in the latest CPUs, a decision is made separately from the others about which instruction to be executed from the queues that hold instructions belonging to all simultaneous threads.
- gpderetta 2y agoAs far as I know hypertreads on intel share fetchers and decoders, so each clock cycle only one thread is feeding in instructions to the pipeline. That's no different to even for a simple barrel processor. It is true that once fetched an OoO CPU does a significant amount of scheduling and it is possible that in a given clock cycle instructions from both threads are getting fed to an execution unit. But I don't think that's the essence of SMT. For example the original larrabee is described as 4-way SMT, but as P5-derived it was a simple in-order design with very limited superscalar capabilities. I very much doubt that at any time instructions from more than one thread were at the execution stage.
- adrian_b 2y agoThe fetchers and decoders are shared, but they fetch and decode many instructions per clock cycle (up to 8 or 9 instructions per clock cycle in the latest Intel cores, i.e. Lion Cove and Skymont). While the shared Intel decoders alternate between the threads and the queue that stores micro-operations before they are dispatched is also partitioned between threads, this front-end is decoupled from the schedulers that select micro-operations for execution, which may choose in any clock cycle as many uops as there are execution units and in any combination between the SMT threads. Even in the first Intel CPU with SMT, Pentium 4, up to 3 instructions were fetched and decoded in each clock cycle and there were places where up to 7 instructions in any combination between the 2 SMT threads were executed during the same clock cycle. In modern CPUs the concurrency is much greater.
- gpderetta 2y agoThey fetch and decode many instructions, but only from a single cache line fetched by the fetcher hence they can't decide instructions for more than one hyperthread at a time. Except the newer *mont cores that have truly separate decoders and fetchers and could indeed decode for two hypertreads separately.