3 ms·
(author here) The FPU is not quite equivalent to a Sandy Bridge FPU, but the FPU is one of the strongest parts of the Bulldozer core. Also, iirc multiply throug
by clamchowder 4y ago
(author here) The FPU is not quite equivalent to a Sandy Bridge FPU, but the FPU is one of the strongest parts of the Bulldozer core. Also, iirc multiply throughput is the same on Bulldozer and K10 at 1 per cycle. K10 uses two pipes to handle high precision multiply instructions that write to two registers, possibly because each pipe only has one result bus and writing two regs requires using both pipes's write ports. But that doesn't mean two multiplies can complete in a single cycle.
With regard to expectations, I don't think AMD ever said that per-thread performance would be a match for Sandy Bridge. ST performance imo was Bulldozer's biggest problem. Calling it 8 cores or 4 cores does not change that. You could make a CPU with say, eight Jaguar cores, and market it as an eight core CPU. It would get crushed by 4c/8t Zen despite the core count difference and no sharing of core resources between threads (on Jaguar).
- adrian_b 4y agoNo, that is not correct. All AMD CPUs since the first Opteron in 2003 until before Bulldozer had a 64-bit integer multiplication throughput of 1 per 2 clock cycles. Initially the Intel CPUs had a much lower throughput, but they improved in each generation, until they matched AMD in Nehalem. In Sandy Bridge, Intel doubled the 64-bit integer multiplication throughput to 1 at each clock cycle. On the other hand, AMD reduced in Bulldozer the 64-bit integer multiplication throughput to 1 per 4 clock cycles. The FPU of Bulldozer had the additional advantage of implementing FMA, but the total throughput in FP multiplications + additions of the 4 Bulldozer FPUs was equal to the total throughput of the 4 Sandy Bridge FPUs. While you are right that calling Bulldozer a 4-core CPU does not change the user expectations about ST performance, it totally changes the user expectations about MT performance. In 2011, the people were not as shocked about the low ST performance (though they were surprised that it was lower than in the AMD Barcelona derivatives), because that was a given ever since Intel had introduced Core 2, as they were shocked about the low MT performance, seeing that an "8-core" CPU is trounced by a 4-core CPU. After being exposed to the AMD propaganda, it was expected that Bulldozer was unlikely to match Intel in ST performance, but it should have a consistent advantage in MT performance, due to being "8-core".
- BlueTemplar 4y agoHuh, so the benchmarks that show that a FX-8xxx is about on par to i3 for ST, but even better than a i5 for MT are misleading ?? I actually have no complaints about MT performance (even though learning that it wasn't a "real" 8-core was disappointing), I can even run VR because it takes full advantage of MT, even though ST is often struggling for other tasks...
- adrian_b 4y agoThe benchmarks that you have in mind had probably been run on some later Piledriver/Steamroller/Excavator models, which corrected some of the initial problems, like the too low instruction decoding throughput, and which raised the clock frequency, due to an improved CMOS SOS process. I also had an AMD Richland APU of 4.4 GHz, which was reasonably faster for most tasks than an Intel Haswell U i5, but the speed ratio was much, much less than the power consumption ratio of 100 W for AMD vs. 15 W for Intel. The original Bulldozer fared much worse against Sandy Bridge.
- BlueTemplar 4y agoHmm, but AFAIK the performance didn't radically improve in P/S/E - and anyway, I had assumed that this whole discussion also covered them because B/P/S/E are still all using the same architecture - for instance : doesn't TFA apply to P/S/E ? (I have one of the last models, IIRC Excavator ?)
- clamchowder 4y agoYeah you're right about the multiplication performance. I checked back and 64-bit integer multiplication is one per four clocks. I disagree that core count should be taken to mean anything about MT performance. You always have to consider the strength of each core too. Nor does twice as many cores for the same architecture imply 2x performance, because there are always shared things like cache and memory bandwidth. And even if those aren't limiting factors, MT boost clocks are often lower than ST ones.