14 ms·
Documentation for the AMD 7900XTX
- secondcoming 3y agoI watched a bit of his attempts to reverse engineer the AMD GPU stack on YT, but Jesus is his keyboard loud.
- pyinstallwoes 3y agodid you learn anything?
- juitpykyk 3y agoI learned that AMD GPUs have layers upon layers of drivers, in user space, kernel space, and drivers running on the device itself.
- brokenmachine 3y agoAnd they didn't really open source anything. I thought they were better than Nvidia in that regard, but they're not. They only open sourced the API. It's still closed firmware just like Nvidia.
- asylteltine 3y ago[dead]
- jpeggtulsa 3y agoIndustry defining architecture?!?!
- tekno45 3y agoCan anyone explain what this level of documentation allows? Could someone make a CUDA compatible tool with this?
- wmf 3y agoYeah, HIP and ZLUDA already exist. Crucially, what the documentation doesn't allow is fixing bugs in the firmware/microcode (all the .bin files).
- deleted 3y ago[deleted]
- roenxi 3y agoSo for reference, this geohot is George Hotz who has a company Tiny Corp [0] in the space. Among a long list of things, he had a moment in the spotlight recently for a long rant [1] where he "gave up on AMD" because their drivers sucked more than he expected. The situation is reasonably complex - AMD have some programmers working on ROCm who seemed to be operating at standard fare when that is really not what AMD needs right now (I personally, suspect there is/was a PM in a key position who didn't "get it", although I am not sure what it is either). As far as I know it got the attention of Lisa Su. I'm cheering him on, even though his complaints were a little melodramatic. My experience is the driver technically supports everything I could possibly want. The problem is if I spend an evening trying to do anything with OpenCL or ROCm the kernel hard-locks and I go to bed early. If the problem inside is what it looks like from the outside (repeating myself, a key manager somewhere just doesn't get the space) they really need some pressure from grumpy customers like George to realign their software development process. That context might be related to this particular case. I see a suspicious folder called "crash" in this repo. [0] https://tinygrad.org/ https://tinygrad.org/ [1] https://www.youtube.com/watch?v=Mr0rWJhv9jU https://www.youtube.com/watch?v=Mr0rWJhv9jU - as I recall he ran the demo suite and it crashed.
- mobilio 3y agoAnd follow-up of that video: https://geohot.github.io/blog/jekyll/update/2023/06/07/a-dive-into-amds-drivers.html https://geohot.github.io/blog/jekyll/update/2023/06/07/a-div...
- echelon 3y agoAnyone working on getting us out of the Nvidia monopoly on commerical AI should be praised. We need lots of competition in this space. I'm hoping Google spins off its TPU division and sells chips to everyone that wants them. I'm hoping AMD gets their act together and that they open up their drivers and stack. I'm hoping a lot of other hardware companies and compute stacks start making inroads on this. Nvidia is not just bilking the whole world on margin, but they're also a limiting reagent in the grand scheme of progress.
- sorenjan 3y ago
- snvzz 3y agoThis is quite crude on what it is. e.g. how does it differ from gpuopen[0] ISA documentation released by AMD themselves? 0. https://gpuopen.com/amd-isa-documentation/ https://gpuopen.com/amd-isa-documentation/
- twothreeone 3y agoThe RDNA docs are just the tip of the iceberg.. in order to get to the compute engines (what you actually care about) you have to go through 3 layers of user-space cruft, a kernel driver, and then a firmware layer running half a dozen separate components of the GPU that "manage" the RDNA compute engines. Apparently, these components run on at least 3 different ISAs (some ARM, some F32, and some on RS64) where 2 of them aren't really documented at all. All of this is what he's trying to bypass and talk to the RDNA cores as directly as possible. Modern consumer GPUs are complex beasts.
- jamesy0ung 3y agoRDNA 3 has a platform security processor? Sounds like a stupid idea.
- jsheard 3y agoMaybe that's where they run the HDCP/DRM stuff? That's the most obvious application of a secure enclave on a GPU that I can think of, and would also explain why they won't (or can't) open it up.
- numpad0 3y agoGPU is itself a standalone computer now, so are CPUs, might as well use the same thing for housekeeping tasks like controlling voltages and clocks.
- corey_moncure 3y agoWe're talking about the company that couldn't fix a mouse cursor corruption bug affecting all of their GPU lines for what, 15, 20 years?
- glimshe 3y agoI used to be an AMD diehard fan. Bought their processors and GPUs for close to 20 years until I simply couldn't take it anymore - 5 years ago I started buying NVIDIA GPUs with Intel CPUs and honestly never looked back. Sure I pay more, but it's worth paying more for reliability and software support. Anything else and I'm losing money out of ideology to give to a business that doesn't have its act together.
- deleted 3y ago[deleted]
- shard972 3y agoAs someone who basicially grew up with Intel/Nvidia but now running AMD/AMD, while i could agree with you on the GPU side that AMD still has more work to do, I would hard disagree on the CPU side. Especially with the X3D line, AMD cpus can absolute smoke intel and at worst, come within a single percentage of performance in usually single thread bound scenarios.
- earthling8118 3y agoIt's always intriguing to see other people's takes. I'm in nearly the complete opposite boat: in recent years I've switched to AMD and it feels like all of my hardware problems have gone away.
- Zambyte 3y agoNew hardware is better than old hardware
- Symmetry 3y agoI'm assuming GP runs Windows and you run Linux. NVidia's proprietary drivers are known to be a lot more stable than AMD's. But on the other hand the open source driver and software stack for AMD (partially shared with Intel) is much more stable than NVidia's on Linux.
- llm_trw 3y agoHeaven is using AMD cards for graphics and NVidia for computation. Hell is the reverse. I have an extra Radeon VII in my ML work station because I don't have to fight nvidia graphics drivers to get it to work. I have the NVidia cards so I don't have to fight AMD drivers to get ML drivers to work.
- shmerl 3y agoOh, this is a very neat tip: umr --list-blocks Getting versions of all these IP blocks for a specific card can be confusing. UMR repo for the reference: https://gitlab.freedesktop.org/tomstdenis/umr https://gitlab.freedesktop.org/tomstdenis/umr
- mkj 3y agoOh, Tom St Denis (of libtomcrypt) is working for AMD? That's promising.
- Apisit8473 3y ago[flagged]
- lofaszvanitt 3y agoSo, why AMD acts like they don't care/can't care about this? Is this such a specialized area that the programmers are hard to find or fished away by Nvidia or something else is going on? If the latter, what?
- Modified3019 3y agoGeorge’s video may have helped precipitate another long post/call to action by Gnif regarding the gpu reset bugs that keep even AMD’s enterprise cards from being reliable enough for VFIO: https://www.reddit.com/r/Amd/comments/1bsjm5a/letter_to_amd_ongoing_amd/ https://www.reddit.com/r/Amd/comments/1bsjm5a/letter_to_amd_...
- treffer 3y agoOne thing everyone should keep in mind about this NVIDIA/AMD battle right now: CUDA has been published for 16 years, it's been a huge push by NVIDIA to do GPGPU computation. I remember seeing it as a new thing in university back then, after the advanced shaders that were only available on NVIDIA. NVIDIA pretty rightful has the lead there, because they worked and invested into it for something like 20 years (you could do pretty advanced shaders on NVIDIA pre-CUDA). It only started to pay off recently, and especially with the AI hype (GPU mining was nice, too). Now everybody is looking at the profits and goes like "OMG, I want a part of that cake!", either by competing (AMD / Intel) or by paying less for the cards (basically everyone else in the AI space). But you have to catch up to 16 years of pretty solid software and ecosystem development. And that's only going to work if you have good enough hardware. NVIDIA did the hard work here. They have earned this lead. I am saying this as someone who would rather not buy NVIDIA. I really wish I can soon throw 1-2 7900XTX into a machine and use it for LLMs without issues. But I would also bet that it takes at least a few more years to catch up, even with the massive global interest.
- paulmd 3y agoYes, this is the counterpoint to the “ 55.58% Net Profit Margin last quarter isn't consistent with a functioning market” thread above. Sure, selling shovels during a gold-rush is very profitable… and nvidia invested a long time in building the best shovels for a lot of years where the net profit of doing so was intensely negative. They built the prospector community up and sponsored the development of geological science that spurred advancement of knowledge and practice - using their products, of course. (The Michelin star model of hardware sales - did you know that Michelin actually makes money from that book!? They’re not doing it because they’re financially disinterested, can you believe that!?!?) Anyway net profit is more like 15% in a normal year. Recently it is actually lower, Ada is already lower margin than pascal for example. It is only this high because nvidia finally struck good - and they spent a lot of effort and money that might never have pqidnoffZ
- nullifidian 3y agoI guess George's 5 hour poking and googling around session, fruitless in terms of the actual results(making compute on AMD's consumer rdna3 chips stable), is newsworthy on HN as of now.