27 ms·
Hacked Nvidia 4090 GPU driver to enable P2P
- arthurcolle 2y agoDoes this work on 4060?
- jsheard 2y agoIt'll be nice while it lasts, until they start locking this down in the firmware instead on future architectures.
- mnau 2y agoSure, but that was something that was always going to happen. So it's better to have it at least for one generation instead of no generation.
- HPsquared 2y agoIs this one of those features that's disabled on consumer cards for market segmentation?
- deleted 2y ago[deleted]
- mvkel 2y agoSort of. An imperfect analogy: a small neighborhood of ~15 houses is under construction. Normally it might have a 200kva transformer sitting at the corner, which provides appropriate power from the grid. But there is a transformer shortage, so the contractor installs a commercial grade 1250kva transformer. It can power many more houses than required, so it's operating way under capacity. One day, a resident decides he wants to start a massive grow farm, and figures out how to activate that extra transformer capacity just for his house. That "activation" is what geohot found
- bogwog 2y agoThat's a poor analogy. The feature is built in to the cards that consumers bought, but Nvidia is disabling it via software. That's why a hacked driver can enable it again. The resident in your analogy is just freeloading off the contractor's transformer. Nvidia does this so that customers that need that feature are forced to buy more expensive systems instead of building a solution with the cheaper "consumer-grade" cards targeted at gamers and enthusiasts.
- bpye 2y agoThis isn’t even the first time a hacked driver has been used to unlock some HW feature - https://github.com/DualCoder/vgpu_unlock https://github.com/DualCoder/vgpu_unlock
- captcanuk 2y agoThere was also this https://hackaday.com/2013/03/18/hack-removes-firmware-crippling-from-nvidia-graphics-card/ https://hackaday.com/2013/03/18/hack-removes-firmware-crippl... using resistors and a different one before that used a graphene lead pencil to enable functionality.
- segfaultbuserr 2y agoExcept that in the computer hardware world, the 1250 kVA transformer was used not because of shortage, but because of the fact that making a 1250 kVA transformer on the existing production line and selling it as 200 kVA, is cheaper than creating a new production line separately for making 200 kVA transformers.
- m3kw9 2y agoWhere is the hack in this analogy
- pixl97 2y agoTaking off the users panel on the side of their house and flipping it to 'lots of power' when that option had previously been covered up by the panel interface.
- rustcleaner 2y agoI am sure many will disagree-vote me, but I want to see this practice in consumer devices either banned or very heavily taxed.
- xandrius 2y agoYou're right. Especially because you didn't present your reasons.
- yogorenapan 2y agoCurious as to your reasoning,
- wmf 2y agoOf course power users want an end to price discrimination because it benefits them... at a cost of more expensive products for the masses.
- imtringued 2y agoWell, they have zero incentives to implement and test this feature for consumer GPUs. Multi GPU setups never really worked that well for gaming.
- llm_trw 2y agoSkimming the readme this is p2p over PCIe and not NVLink in case anyone was wondering.
- klohto 2y agoafaik 4090 doesn’t support 5.0 so you are limited to 4.0 speeds. Still an improvement.
- formerly_proven 2y agoRTX 40 doesn’t have NVLink on the PCBs, though the silicon has to have it, since some sibling cards support it. I’d expect it to be fused off.
- deleted 2y ago[deleted]
- HeatrayEnjoyer 2y agoHow to unfuse it?
- magicalhippo 2y agoI don't know about this particular scenario, but typically fuses are small wires or resistors that are overloaded so they irreversibly break the connection. Hence the name. Either done during manufacture or as a one-time programming[1][2]. Though sometimes reprogrammable configuration bits are sometimes also called fuse bits. The Atmega328P of Arduino fame uses flash[3] for its "fuses". [1]: https://www.nxp.com/docs/en/application-note/AN4536.pdf https://www.nxp.com/docs/en/application-note/AN4536.pdf [2] https://www.intel.com/programmable/technical-pdfs/654254.pdf https://www.intel.com/programmable/technical-pdfs/654254.pdf [3]: https://ww1.microchip.com/downloads/en/DeviceDoc/Atmel-7810-Automotive-Microcontrollers-ATmega328P_Datasheet.pdf https://ww1.microchip.com/downloads/en/DeviceDoc/Atmel-7810-...
- HeatrayEnjoyer 2y ago
- awsanswers 2y ago[flagged]
- jagrsw 2y agoWas it George himself, or a person working for a bounty that was set up by tinycorp? Also, a question for those knowledgeable about the PCI subsys: it looked like something NVIDIA didn't care about, rather than something they actively wanted to prevent, no?
- mtlynch 2y agoCommits are by geohot, so it looks like George himself.
- throw101010 2y agoI've seen him work on tinygrad on his Twitch livestream couple times, so more than likely him indeed.
- deleted 2y ago[deleted]
- squarra 2y agoHe also documented his progress on the tinygrad discord
- throwaway8481 2y agoI feel like I should say something about discord not being a suitable replacement for a forum or bugtracker.
- guywhocodes 2y agoWe are talking about a literal monologue while poking at a driver for a few hours, this wasn't a huge project.
- toast0 2y agoPCI devices have always been able to read and write to the shared address space (subject to IOMMU); most frequently used for DMA to system RAM, but not limited to it. So, poking around to configure the device to put the whole VRAM in the address space is reasonable, subject to support for resizable BAR or just having a fixed size large enough BAR. And telling one card to read/write from an address that happens to be mapped to a different card's VRAM is also reasonable. I'd be interested to know if PCI-e switching capacity will be a bottleneck, or if it'll just be the point to point links and VRAM that bottlenecks. Saving a bounce through system RAM should help in either case though.
- klohto 2y agofyi should work on most 40xx[1] [1] https://github.com/pytorch/pytorch/issues/119638#issuecomment-2051196015 https://github.com/pytorch/pytorch/issues/119638#issuecommen...
- clbrmbr 2y agoIf we end up with a compute governance model of AI control [1], this sort of thing could get your door kicked in by the CEA (Compute Enforcement Agency). [1] https://podcasts.apple.com/us/podcast/ai-safety-fundamentals-alignment/id1680794263?i=1000651665081 https://podcasts.apple.com/us/podcast/ai-safety-fundamentals...
- logicchains 2y agoLooks like we're only a few years away from a bona fide cyberpunk dystopia, in which only governments and megacorps are allowed to use AI, and hackers working on their own hardware face regular raids from the authorities.
- tomoyoirl 2y agoMere raids from the authorities? I thought EliY was out there proposing airstrikes.
- the8472 2y agoIn the sense that any other government regulation is also ultimately backed by the state's monopoly on legal use of force when other measures have failed. And contrary to what some people are implying he also proposes that everyone is subject to the same limitations, big players just like individuals. Because the big players haven't shown much of a sign of doing enough.
- tomoyoirl 2y ago> In the sense that any other government regulation is also ultimately backed by the state's monopoly on legal use of force when other measures have failed. Good point. He was only (“only”) really calling for international cooperation and literal air strikes against big datacenters that weren’t cooperating. This would presumably be more of a no-knock raid, breaching your door with a battering ram and throwing tear gas at the wee hours of the morning ;) or maybe a small extraterritorial drone through your window
- ewalk153 2y agoDoes this appear to be intentionally left out by NVidia or an oversight?
- creshal 2y agoSeems more like an oversight, since you have to stitch together a bunch of suboptimal non-default options?
- arghwhat 2y agoIt does seem like an oversight, but there's nothing "suboptimal non-default options" about iteven if the implementation posted here seems somewhat hastily hacked together.
- segfaultbuserr 2y ago> but there's nothing "suboptimal non-default options" about it If "bypassing the official driver to invoke the underlying hardware feature directly through source code modification (and incompatibilities must be carefully worked around by turning off IOMMU and large BAR, since the feature was never officially supported)" does not count as "suboptimal non-default options", then I don't know what counts as "suboptimal non-default options".
- talldayo 2y ago> then I don't know what counts as "suboptimal non-default options". Boy oh boy do I have a bridge to sell you: https://nouveau.freedesktop.org/ https://nouveau.freedesktop.org/
- _zoltan_ 2y agoI have some news for you: you must disable IOMMU on the H100 platform anyway, at least for optimal GDS :-)
- segfaultbuserr 2y ago
- casiejudi 2y ago[dead]
- rfoo 2y agoGlad to see that geohot is back being geohot, first by dropping a local DoS for AMD cards, then this. Much more interesting :p
- deleted 2y ago[deleted]
- jaimehrubiks 2y agoIs this the same guy that hacked the PS3?
- mepian 2y agoYes, that's him.
- deleted 2y ago[deleted]
- WithinReason 2y agoAnd the iPhone
- yrds96 2y agoAnd android
- zoklet-enjoyer 2y agoAnd the crypto scam cheapETH
- mattanimation 2y ago[flagged]
- deleted 2y ago[deleted]
- gigatexal 2y agoas a technical feat this is really cool! though as others mention i hope you don't get into too much hot water legally seems anything that remotely lets "consumer" cards canibalize anything with the higher end H/A-series cards Nvidia would not be fond of and they've the laywers to throw at such a thing
- jstanley 2y agoWhat does P2P mean in this context? I Googled it and it sounds like it means "peer to peer", but what does that mean in the context of a graphics card?
- deleted 2y ago[deleted]
- haunter 2y agoShared memory access for Nvidia GPUs https://developer.nvidia.com/gpudirect https://developer.nvidia.com/gpudirect
- deleted 2y ago[deleted]
- __alexs 2y agoIt means you can send data from the memory of 1 GPU to another GPU without going via RAM. https://xilinx.github.io/XRT/master/html/p2p.html https://xilinx.github.io/XRT/master/html/p2p.html
- ot1138 2y agoIs this really efficient or practical? My understanding is that the latency required to copy memory from CPU or RAM to GPU negates any performance benefits (much less running over a network!)
- brrrrrm 2y agoYea. It’s one less hop through slow memory
- whereismyacc 2y agothis would be directly over the memory bus right? I think it's just always going to be faster like this if you can do it?
- 2y ago
- ivanjermakov 2y agoI was always fascinated by George Hotz's hacking abilities. Inspired me a lot for my personal projects.
- vrnvu 2y agoI agree, I feel so inspired with his streams. Focus and hard work, the key to good results. Add a clear vision and strategy, and you can also accomplish “success”. Congratulations to him and all the tinygrad/comma contributors.
- sambull 2y agoHe's got that focus like a military pilot on a long flight.
- Jerrrry 2y agoHis Xbox360 laptop was the crux of teenage-motivation, for me.
- jgpc 2y agoI agree. It is fascinating. When you observe his development process (btw, it is worth noting his generosity in sharing it like he does) he gets frequently stuck on random shallow problems which a perhaps more knowledgable engineer would find less difficult. It is frequent to see him writing really bad code, or even wrong code. The whole twitter chapter is a good example. Yet, himself, alone just iterating resiliently, just as frequently creates remarkable improvements. A good example to learn from. Thank you geohot.
- namibj 2y agoAnd here I thought (PCIe) P2P was there since SLI dropped the bridge (for the unfamiliar, it looks and acts pretty much like an NVLink bridge for regular PCIe slot cards that have NVLink, and was used back in the day to share framebuffer and similar in high-end gaming setups).
- wmf 2y agoSLI was dropped years ago so there's no need for gaming cards to communicate at all.
- userbinator 2y agoI wish more hardware companies would publish more documentation and let the community figure out the rest, sort of like what happened to the original IBM VGA (look up "Mode X" and the other non-BIOS modes the hardware is actually capable of - even 800x600x16!) Sadly it seems the majority of them would rather tightly control every aspect of their products' usage since they can then milk the userbase for more $$$, but IMHO the most productive era of the PC was also when it was the most open.
- rplnt 2y agoThen they couldn't charge different customers different amounts for the same HW. It's not a win for everyone.
- axus 2y agoThe price of 4090 may increase now, in theory locking out some features might have been a favor for some of the customers.
- Sayrus 2y agoBut it wouldn't if all cards supporting this were "unlocked" by default and thus the other "enterprise-grade" cards weren't that much more expensive. Of course that'd reduce profits by a lot.
- paulmd 2y agoit probably would - you saw exactly that outcome with mining. for a lot of these demand bursts, demand is so high it cannot be sated even consuming 100% or 200% of typical GPU production. cards like RX 6500XT that simply don't have the RAM to participate were less affected, but even then you've got enough cross-elasticity (demand from people being crowded out of other product segments) that tends to pump prices to 2-3x the "normal" clearance prices we see today. And yes, absolutely anything that can mine in any capacity will get pulled in during that sort of boom/bubble, not just "high-end"/"enterprise".
- 2y ago
- andersa 2y agoIncredible! I'd been wondering if this was possible. Now the only thing standing in the way of my 4x4090 rig for local LLMs is finding time to build it. With tensor parallelism, this will be both massively cheaper and faster for inference than a H100 SXM. I still don't understand why they went with 6 GPUs for the tinybox. Many things will only function well with 4 or 8 GPUs. It seems like the worst of both worlds now (use 4 GPUs but pay for 6 GPUs, don't have 8 GPUs).
- corn13read2 2y agoA macbook is cheaper though
- tgtweak 2y agoThe extra $3k you'd spend on a quad-4090 rig vs the top mbp... ignoring the fact you can't put the two on even ground for versatility (very few libraries are adapted to apple silicone let alone optimized). Very few people that would consider an H100/A100/A800 are going to be cross-shopping a macbook pro for their workloads.
- LoganDark 2y ago> very few libraries are adapted to apple silicone let alone optimized This is a joke, right? Have you been anywhere in the LLM ecosystem for the past year or so? I'm constantly hearing about new ways in which ASi outperforms traditional platforms, and new projects that are optimized for ASi. Such as, for instance, llama.cpp.
- cavisne 2y agoNothing compared to Nvidia though. The FLOPS and memory bandwidth is simply not there.
- spudlyo 2y agoThe memory bandwidth of the M2 Ultra is around 800GB/s verses 1008 GB/s for the 4090. While it’s true the M2 has neither the bandwidth or the GPU power, it is not limited to 24G of VRAM per card. The 192G upper limit on the M2 Ultra will have a much easier time running inference on a 70+ billion parameter model, if that is your aim. Besides size, heat, fan noise, and not having to build it yourself, this is the only area where Apple Silicon might have advantage over a homemade 4090 rig.
- xipho 2y agoYou can watch this happen on the weekends, typically, sometimes, for some very long sessions, sometimes. https://www.twitch.tv/georgehotz https://www.twitch.tv/georgehotz
- BeefySwain 2y agoCan someone ELI5 what this may make possible that wasn't possible before? Does this mean I can buy a handful of 4090s and use it in lieu of an h100? Just adding the memory together?
- segfaultbuserr 2y agoNo. The Nvidia A100 has a multi-lane NVLink interface with a total bandwidth of 600 GB/s. The "unlocked" Nvidia RTX 4090 uses PCIe P2P at 50 GB/s. It's not going to replace A100 GPUs for serious production work, but it does unlock a datacenter-exclusive feature and has some small-scale applications.
- xmorse 2y agoFinally switched to Nvidia and already adding great value
- perfobotto 2y agoWhat stops nvidia from making sure this stops working in future driver releases?
- __MatrixMan__ 2y agoThe law, hopefully. Beeper mini only worked with iMessage for a few days before Apple killed it. A few months later the DOJ sued Apple. Hacks like this show us the world we could be living in, a world which can be hard to envision otherwise. If we want to actually live in that world, we have to fight for it (and protect the hackers besides).
- StayTrue 2y agoI was thinking the same but in terms of firmware updates.
- aresant 2y agoSo assuming you utilized this with (4) x 4090s is there a theoretical comparative to performance vs the A6000 / other professional lines?
- thangngoc89 2y agoI believe this is mostly for memory capacities. PCIe access between GPUs is slower than soldered RAM on a single GPU
- andersa 2y agoIt depends on what you do with it and how much bandwidth it needs between the cards. For LLM inference with tensor parallelism (usually limited by VRAM read bandwidth, but little exchange needed) 2x 4090 will massively outperform a single A6000. For training, not so much.
- c0g 2y agoAny idea of DDP perf?
- No1 2y agoThe original justification that Nvidia gave for removing Nvlink from the consumer grade lineup was that PCIe 5 would be fast enough. They then went on to release the 40xx series without PCIe 5 and P2P support. Good to see at least half of the equation being completed for them, but I can’t imagine they’ll allow this in the next gen firmware.
- musha68k 2y agoOK now we are seemingly getting somewhere. I can feel the enthusiasm coming back to me. Especially in light of what's going on with LocalLLaMA etc: https://www.reddit.com/r/LocalLLaMA/comments/1c0mkk9/mistral_8x22b_already_runs_on_m2_ultra_192gb_with https://www.reddit.com/r/LocalLLaMA/comments/1c0mkk9/mistral...
- thangngoc89 2y ago> You may need to uninstall the driver from DKMS. Your system needs large BAR support and IOMMU off. Can someone point me to the correct tutorial on how to do these things?
- unaindz 2y agoThe first one I assume is the nvidia driver for linux installed using dkms. If it uses dkms or not is stated on the drivers name, at least on arch based distributions. The latter options are settings on your motherboard bios, if your computer is modern, explore your bios and you will find them
- jasomill 2y agoDKMS: uninstall Nvidia driver using distro package manager BAR: enable resizable BAR in motherboard CMOS setup IOMMU: Add "amd_iommu=off" or "intel_iommu=off" to kernel command line for AMD or Intel CPU, respectively (or just add both). You may or may not need to disable the IOMMU in CMOS setup (Intel calls its IOMMU VT-d). See motherboard docs for specific option names. See distro docs for procedures to list/uninstall packages and to add kernel command line options.
- spxneo 2y agodoes this mean you can horizontally scale to GPT-4-esque LLM locally in the near future? (i hear you need 1TB of VRAM) Is Apple's large VRAM offering like 196gb offer the fastest bandwidth and if so how will pairing a bunch of 4090s like in the comments work?
- lawlessone 2y agoThis is very interesting. I can't afford two mortgages though ,so for me it will have to just stay as something interesting :)
- m3kw9 2y agoIn layman terms what does this enable?
- vladgur 2y agocurious if this will ever make it to 3090s
- cavisne 2y agoHow does this compare in bandwidth and latency to nvlink? (I’m aware it’s not available on the consumer cards)
- modeless 2y agoWhat are the chances that Nvidia updates the firmware to disable this and prevents downgrading with efuses? Someday cards that still have older firmware may be more valuable. I'd be cautious upgrading drivers for a while.
- theturtle32 2y agoWTF is P2P?
- theturtle32 2y agoAnswered my own question with a Google search: https://developer.nvidia.com/gpudirect#:~:text=LEARN%20MORE%20%E2%80%BA-,GPUDirect%20Peer%20to%20Peer,fabric%20(PCIe%2C%20NVLink) https://developer.nvidia.com/gpudirect#:~:text=LEARN%20MORE%.... > GPUDirect Peer to Peer > Enables GPU-to-GPU copies as well as loads and stores directly over the memory fabric (PCIe, NVLink). GPUDirect Peer to Peer is supported natively by the CUDA Driver. Developers should use the latest CUDA Toolkit and drivers on a system with two or more compatible devices.
- deleted 2y ago[deleted]
- tanelpoder 2y agoI also love that it can be done with just a few code line changes: https://github.com/NVIDIA/open-gpu-kernel-modules/commit/1f4613dacec2638569a74b5e3dbcab01832f72a7?diff=unified&w=1 https://github.com/NVIDIA/open-gpu-kernel-modules/commit/1f4...
- waldrews 2y agoWould this approach be possible to extend downmarket, to older consumer cards? For a lot of LLM use cases we're constrained by memory and can tolerate lower compute speeds so long as there's no swapping. ELI5, what would prevent a hundred 1060-level cards from being used together?
- Sebb767 2y ago> ELI5, what would prevent a hundred 1060-level cards from being used together? In this case, you'd only have a single PCIe (v3!) lane per card, making the interconnect speed horribly slow. You'd also need to invest in thousands of dollars of hardware to get all of those cards connected and, unless power is free, you'd outspend any theoretical savings instantly. In general, if you go back in card generations, you'll quickly hit so low memory limits amd slower compute that a modern CPU-based setup is better value.
- qxfys 2y agoI am amazed how people always find a way to make this kind of thing work. kudos!
- chriskanan 2y agoThis is great news. As an academic, I'm aware of multiple labs that built boxes with 4090s, not realizing that Nvidia had impaired P2P communication among cards. It's one of the reasons I didn't buy 4090s, despite them being much more affordable for my work. It isn't nvlink, but Nvidia has mostly gotten rid of that except for their highest end cards. It is better than nothing. Late last year, I got quotes for machines with four nvlink H100s, but the lead time for delivery was 13 months. I could get the non-nvlink ones in just four months. For now, I've gone with four L40S cards to hold my lab over but supply chain issues and gigantic price increases are making it very hard for my lab to do it's work. That's not nearly enough to support 6 PhD students and a bunch of undergrads. Things were a lot easier when I could just build machines with two GPUs each with Nvlink for $5K each and give one to each student to put under their desks, which is what I did back in 2015-2018 at my old university.
- uniqueuid 2y agoAnd before that, Nvidia made our lives harder by phasing out blower-style designs in consumer cards that we could put in servers. In my lab, I'd take a card for 1/4 the price that has half the MTBF over a card for full price anytime.
- photonbeam 2y agoHow does cost compare with some of the GPU-cloud providers?
- uniqueuid 2y agoNot op, but I found this benchmark of whisper large-v3 interesting [1]. It includes the cloud provider's pricing per gpu, so you can directly calculate break-even timing. Of course, if you use different models, training, fine tuning etc. the benchmarks will differ depending on ram, support of fp8 etc. [1] https://blog.salad.com/whisper-large-v3/ https://blog.salad.com/whisper-large-v3/
- jeffs4271 2y agoIt is cool seeing hacks like this. But this is something to be careful with, as GH100 had hardware changes to meet CUDA fence requirements.
- gururise 2y agoHow long before Nvidia patches this?
- lucifer_is_back 2y agoso basically rtx 4090 x6 = 144 GB ram which would cost $15996 = $9594 ( only the nvidia 4090s) and currently the tiny box gives *TinyBox* > GPU RAM | 144 GB > Price | $15,000 $25,000 Nvidia 4090x6 > GPU RAM | 144 GB > Price | $9594 so a * 36.04%* decrease in price from team red tinybox ( $15k) and *61.624% *decrease in price from the team green tinybox ( $25k)