10 ms·
Reverse engineering Dell iDRAC to get rid of GPU throttling
- Terminal135 3y agoThe repo claims that the servers themselves throttle the GPUs, but isn't it the GPUs themselves that can throttle or maybe the OS? Neither of those are controlled by the server (hopefully) so is there a different system at play here?
- csdvrx 3y agoNo, that's controlled by the server: try lspci -vv on any linux system. Look at the link speed and width, like LnkSta: Speed 8GT/s, Width x2: x2 means 2 lanes. Try: `sudo lspci -vv | grep -P "[0-9a-f]{2}:[0-9a-f]{2}\.[0-9a-f]|downgrad" |grep -B1 downgrad` Besides the speed, you can have another problem with lanes limitations. For example, AMD CPUs have a lot of lanes, but unless you have an EPYC, most of them are not exposed, so the PCH tries to spread its meager set among the devices connected to your PCI bus, and if you have a x16 GPU, but also a WIFI adapter, a WWAN card and a few identical NVMe, you may find only of the NVMe benchmarks at the throughput you expect.
- toast0 3y ago> For example, AMD CPUs have a lot of lanes, but unless you have an EPYC, most of them are not exposed, so the PCH tries to spread its meager set among the devices connected to your PCI bus, and if you have a x16 GPU, but also a WIFI adapter, a WWAN card and a few identical NVMe, you may find only of the NVMe benchmarks at the throughput you expect. Most AM4 boards put an x16 slot direct to the CPU, and an x4 direct linked NVMe slot. That's 20 of the 24 lanes; the other 4 lanes go to the chipset, which all the rest of the peripherals are behind. (There's some USB and other I/O from the cpu, too). AM5 CPUs added another 4 lanes, which is usually a second cpu x4 slot. Early AM4 boards might not have a cpu x4 NVMe slot, and those 4 cpu lanes might not be exposed, and the a300/x300 chipsetless boards don't tend to expose everything, but where else are you seeing AMD boards where all the CPU lanes aren't exposed?
- csdvrx 3y ago> Early AM4 boards might not have a cpu x4 NVMe slot, and those 4 cpu lanes might not be exposed, and the a300/x300 chipsetless boards don't tend to expose everything I'm sorry, I oversimplified, and said "most of them" while I should have said "not all of them" as 20/24 is more correct for B550 chipsets (the most common for AM4) instead of trying to generalize. Your explanation is more correct that mine. For anyone who might want extra details about the number of lanes per CPU, https://pcguide101.com/motherboard/how-many-pcie-lanes-does-ryzen-have/ https://pcguide101.com/motherboard/how-many-pcie-lanes-does-... is a good read that shows the difference for APUs.
- toast0 3y agoI'm still not quite sure what you're trying to say? Lanes behind the chipset are multiplexed, and you can't get more than x4 throughput through the chipset (and the link speed between the cpu and the chipset varies depending on the chipset and cpu). But that's not a problem of the CPU lanes not being exposed, it's a problem of "not enough lanes" or more likely, lanes not arranged how you'd like. On AM4, if your GPU uses x16, and one NVMe uses x4, then everything else is going to be squeezed through the chipset. On AM5, you usually get two x4 NVMe slots, but again everything else is squeezed through the chipset; x670 is particularly constrained because it just puts a second chipset downstream of the first chipset, so you're just adding more stuff to squeeze through the same x4 link to the CPU. Personally, I found that link to be more confusing than just reading through the descriptions on wikipedia for a particular Zen version. For example https://en.wikipedia.org/wiki/Zen_3 https://en.wikipedia.org/wiki/Zen_3 ... just text search in the page for "lanes" and it explains for all the flavors of chips how many lanes, and how many go to the chipset. Similarly the page for AMD chipsets is pretty succinct https://en.wikipedia.org/wiki/List_of_AMD_chipsets#AM5_chipsets https://en.wikipedia.org/wiki/List_of_AMD_chipsets#AM5_chips...
- formerly_proven 3y agoThere's a reason why so many motherboard makers avoid putting a block diagram in their manuals and go for paragraphs of legalese instead, and laziness is only half of it.
- ilyt 3y ago> For example, AMD CPUs have a lot of lanes, but unless you have an EPYC, most of them are not exposed, so the PCH tries to spread its meager set among the devices connected to your PCI bus, and if you have a x16 GPU, but also a WIFI adapter, a WWAN card and a few identical NVMe, you may find only of the NVMe benchmarks at the throughput you expect. example from my X670E board * first NVME = 4x gen 5 * second= 4x gen 4 * 2 USB ports connected to CPU (10/5 Gbit) and EVERYTHING ELSE goes thru 4x gen 4 PCIE bus, including additional 3x nvme, 7 SATA ports, a bunch of USBs, few 1x PCIE ports, network, etc.
- formerly_proven 3y agoPCIe devices can only draw a limited wattage until the host clears them for higher power. There is also a separate power brake mechanism (optional part of PCIe) mentioned in the article, which has been proposed by nVidia for PCIe so it seems likely their GPUs support it.
- f_devd 3y agoI can actually answer this (as it is how I stumbled on to the repo), it's through a signal from the motherboard called Pwrbrk (Power Brake), Pin 30 on PCIe. It tells the PCIe device to maintain a low-power mode, in the case of Nvidia GPUs it's about 50W (300Mhz out of 2100Mhz in my case). You can check if it's active using `nvidia-smi -q | grep Slowdown` as shown in the post
- bubblethink 3y agoDell is peak asshole design. They also blast fans as full speed if you install GPUs that you don't buy from them. Fuck them.
- flykespice 3y agoIsn't that illegal? Hijacking your customer pc if they install something that isn't from them?
- Maxburn 3y agoNominally done for your protection. Lowering power (clock) and heat load (fast fan) for unapproved gear prevents things from going dead and getting people REALLY mad and likely reduces warranty claims.
- somehnguy 3y agoRestrictive nonsense seems common in the server space unfortunately. HPe do similar things. IIRC they disabled certain features if you used non-HPe ‘approved’ hard drives.
- bg46z 3y agoDell also does this with their EMC storage arrays, it’s meant to push you towards their pro services. You are supposed to tell the array to order drives for you from pro services and someone from some nameless MSP contracted with dell installs it for you at a 10x markup.
- formerly_proven 3y agoAre there even any good server vendors? Dell, HPE and Lenovo do their lock-in shit. Supermicro's BMC is pretty bad. xFusion is totally-not-Huawei-I-pwomise. There's a few more that come to mind but all of them are niches like HPC and don't really do sales on a small scale.
- csdvrx 3y agoOther manufacturers do worse and prevent boot if your PCI ids aren't on a positive list. This is for example present on thinkpads, and while you could patch the bios before, Intel bootguard now prevents you do that "for your own protection" :) I hope the MSI leak contains actual bootguard keys for intel 11th gen+, and can be used to allow "unauthorized" PCI modules on modern thinkpads!
- deleted 3y ago[deleted]
- somat 3y agoBMC's in general leave me uneasy. I like the idea, it is a small computer that is used to monitor and control your big computer. But hate the implementation. Why are they all super secret special firmware blobs? Why can't I just install my linux of choice and run the manufacturers software? This would still suck but not as bad as the full stack nonsense they foist on you at this point.
- donalhunt 3y agoWas responsible for trying to improve the management and operation of a large fleet of BMCs for a while. Plenty of bugs and pace of releases is slow. :( Definitely an area where a more open ecosystem would improve the pace of innovation.
- hinkley 3y agoWe need a Linux for BMCs. Oxide is working on one, but I'd like to see a contender fielded from the seL4 community, along with some other folks. For example, why doesn't Wind River have one already?
- NexRebular 3y agoWhy linux? Why not *BSD?
- gaius_baltar 3y ago> Why linux? Why not *BSD? GPL can force manufacturers to cooperate with users. Of course, they can still use closed source binary modules and userland programs ...
- NexRebular 3y agoGPL can also turn manufacturers away. I would rather have variation in the possible BMC operating systems instead of sticking linux everywhere and contributing to a monoculture.
- amir734jj 3y agoI'm dealing with something similar. I wanted to use Redfish to clear out hard drives but storage is not standardize across different vendors. Dell has a secure erase. HPE gen10 has smart storage and anything older doesn't have any useful functionality in their Redfish API. What a mess. So I need to use PXE booting and probably winpe to do this.
- walrus01 3y agoThere are a number of valid engineering reasons for thermal dissipation why you don't want to overload the heat producing things in a 1U server beyond what it was designed for. This article doesn't mention at all what the max TDP of each gpu is, which makes me suspicious. Or things like max tdp of cpus (such as when running a prime number calculatio multi core stress benchmark to load them to 100%) combined with total wattage of GPUs. If you have never built an x86-64 1U dual socket server from discrete whitebox components (chassis, power supply, 12x13 size motherboard, etc) this is harder to intuitively understand. I would recommend that people who want four powerful GPUs in something they own themselves to look at more conventional sized server chassis, 3U to 4U in height, or tower format if it doesn't need to be in a datacenter cabinet somewhere.