5 ms·
Based on how bulldozer performed, it'll end up that sure it's 5ghz, but did we mention all our instructions take two times the number of clocks now? Virtualizt
by trotsky 13y ago
Based on how bulldozer performed, it'll end up that sure it's 5ghz, but did we mention all our instructions take two times the number of clocks now?
Virtualiztion host performance on our 8 core bulldozer (esx 5.1, private kbs from vmware to try to help, 32gb ram, rad10 zfs san) was so bad (think p4 era) that we finally tracked down how to force the cpu into only using 4 cores, one per real fp core.
The reality is that there is no mainstream scheduler out that that can efficienty use cores set up like that, especially with the long pipelines. I'm not sure it can't be done, but what improvements have been made have been minimal, or just in an academic/not a real os situation.
That's why Intel ships a compiler, duh.
It is true that the # one thing holding that part back was the raw clock speed (as long as you view it more like a 4 core, 8 thread part ala Intel), but i've gone back to speccing intel - it's just not worth being that much of a ginae pig for a firm thats basically trying to scrape by until the arm64 parts start getting stamped.
- sliverstorm 13y agowe finally tracked down how to force the cpu into only using 4 cores Was the issue the shared fp units, or the turbo-core? I wonder if you can disable the turbo-core?
- mitchty 13y agoIf it was anything like the Niagra processors the shared fp units normally are a bottleneck for fp. But the larger problem was the register remapping/pipelines. They were fast, if you were running certain workloads. God help you if you had to compress anything on those systems. Without pbzip2 or pigz it took forever. Really bad example but bulldozer seemed way too niagraish to me based on its goals.
- AnthonyMouse 13y agoRunning threaded floating point workloads on bulldozer-derived architectures is just folly. If you have parallel floating point code you should in general be running it on a GPU.
- trotsky 13y agoThese weren't fp intensive workloads at all - mostly your typical IT IO workloads. I don't know the internals to say exactly what or why, but something seems to go really wrong on bulldozer when you try to schedule two different vms on the same coupled pair of cores.
- AnthonyMouse 13y agoIt's because they're not independent cores. You're pretty much never going to get the same single-thread performance with two threads running on a module as with one, the idea is that you ought to get better than 0.5X the single thread performance, such that if you have two threads then 2*0.75X is better than X, while still allowing you to get X (or better with turbo) on strictly single threaded workloads. Where this can fall apart is if you're trying to use eight homogenous threads at once and the threads have large working set sizes, such that the second thread causes spill out from the per-module caches. Then you have eight threads contending for L3 bandwidth, or if you're really screwed you fill up the L3 and start to hit main memory. Out of curiosity, have you tried any of the Abu Dhabi Opterons? They doubled the L3 from 8MB to 2x8MB, which I would expect to help by both keeping you out of main memory and reducing contention by splitting each L3 between half as many cores (assuming you don't get the new twice-as-many-cores models).
- trotsky 13y agoSorry for the late reply. I am just guessing that it was the shared fp, since that's what i thought the major shared resource was. The workloads were bog standard kind of stuff, so i assume mostly integer work, so it's absolutely possible that it wasn't the fp sharing but something else. I thought that the integer cores all had their own registers though? Never the less, I don't know enough about cpu internals to say - I was just assuming.
- kinghajj 13y agoDo you know if it used Bulldozer or the updated Piledriver architecture? I just set up a personal home server yesterday with a 8350/16GB and ESXi 5.1, and the performance seems fine enough. Are there any kinds of tasks where the slowness becomes most apparent?
- sliverstorm 13y agoYour parent was using a SAN, which means your parent was in an enterprise setting and probably running 8 guests minimum (1 per core). I'd hazard that is where issues started to crop up.
- trotsky 13y agosorry for the late reply - that machine was using bulldozer cores, not piledriver. Also I was exaggerating for effect I guess, that thing has left a bad taste in my mouth. What I really mean is that when compared to a mainstream i7 it sucks - you can effectively put twice as much work on the i7, and it has the added bonus of being able to run a cpu hungry single threaded job basically twice as fast as the bulldozer when there isn't much/any cpu contention.