11 ms·
AMD's 7900 XTX achieves better value for Stable Diffusion than Nvidia RTX 4080
- Der_Einzige 3y agoNo it doesn’t. AMD drivers don’t support all of the extensions, optimizations, and related in things like automatic1111. There’s always stuff that breaks on AMD and works perfectly in CUDA land.
- FloatArtifact 3y agoThey explicitly use Automatic1111 to demonstrate those speeds... Granted maybe not all extensions are compatible with DirectML in Automatic1111.
- roenxi 3y agoI do think the original comment was wrong (if they tested in Automatic1111 then ... enough said about it working). However, they still have a point in warning people about AMD cards. There is more risk to these things than Nvidia: - AMD's support is flaky. The AMD 7900 XTX is only officially supported on Windows PCs (for ROCm). On Linux PCs in my experience there is a high risk of graphics card hard lockups. - Historically AMD has dropped support for consumer cards really quickly (I don't think they support any consumer cards from more than about 3 years ago). At some point that'll stabilise, but it is hard to tell if this is the generation that will start working long term or the next. - AMD generally gets second class support in return from the machine learning communities. There still seems to be some chance that you'll be locked out of the latest and greatest if you go AMD. At some point this will all probably come good and it'll be like CPUs where there is really no difference between the vendors. There is too much interest in machine learning right now for anything else to happen ... but I've been buying AMD for more than a decade now and I've been thinking that for years already. If someone is on the edge they'd be better off not getting seduced by a minor performance improvement and sticking to Nvidia cards. It is uncertain when the long term will reach us. This year? 5 years? A decade? All options where AMDs software drivers are in play. Although AMDs graphics drivers do seem great on Linux nowadays, I like things that just work. It is a pity they got so out of position vs CUDA.
- nextaccountic 3y ago> On Linux PCs in my experience there is a high risk of graphics card hard lockups. What do you mean, does this become a crash / kernel panic? Is this a common occurrence and is it due to bad kernel drivers?
- paulmd 3y agoinfamously, geohot encountered and documented a bunch of these kernel panics. https://github.com/RadeonOpenCompute/ROCm/issues/2198 https://github.com/RadeonOpenCompute/ROCm/issues/2198 https://geohot.github.io/blog/jekyll/update/2023/06/07/a-dive-into-amds-drivers.html https://geohot.github.io/blog/jekyll/update/2023/06/07/a-div... they don't occur in amdgpu-pro proprietary userland, but AMD apparently runs that on an internal release cycle and doesn't make the current repo or build available to the public. And AMD just historically hasn't cared enough about the open userland to support that even in the officially-supported configurations. I'll also add to GP by saying that the windows support is all of about 2 weeks old at this point and it's also the first time that consumer cards (as opposed to the CDNA compute cards and the workstation-branded Radeon Pro cards) have had official support. Not that that is a guarantee it'll actually run with AMD though, especially on the open-source userland.
- roenxi 3y agoIt looks like a kernel panic to me, probably bad kernel drivers but I've never tried to figure out what is happening. My graphics card is more than 3 years old (ie, unsupported) so it isn't worth reporting.
- sassy_quat 3y ago[dead]
- Jackson__ 3y agoYeah, if they're going to compare techniques breaking various features and extensions, I think it's only fair to compare it to similarly incompatible techniques for nvidia. In which case, I would recommend the AITemplate extension[0] for ComfyUI, which runing at a (hopefully) standardized 512x512 res, default euler A sampler, nets me about 23.8 it/s on my 350w 3090. [0] https://github.com/FizzleDorf/AIT https://github.com/FizzleDorf/AIT
- dragonwriter 3y ago> AMD drivers don’t support all of the extensions, optimizations, and related in things like automatic1111. With olive in general, you need to rebuild all the models for it; and a whole lot of the extensions for A1111 are support for workflow components with new models (often, whole new classes of models). Unless Olive becomes a major platform for people developing models, this is always going to be lagging.
- smoldesu 3y agoWait, why are they comparing Microsoft Olive on AMD to Pytorch on Nvidia? Nvidia supposedly shipped support for Olive recently, there should be no problem getting a head-to-head comparison: https://www.tomshardware.com/news/nvidia-geforce-driver-promises-doubled-stable-diffusion-performance https://www.tomshardware.com/news/nvidia-geforce-driver-prom... This is a very strange comparison.
- asu_thomas 3y agoA head-to-head comparison would render less ad views.
- lostmsu 3y agoMy understanding is Olive is a compressor, so comparing olive results to raw model is invalid.
- dragonwriter 3y ago> Nvidia supposedly shipped support for Olive recently I mean, they announced it with a more than 2x speedup for SD in May: https://blogs.nvidia.com/blog/2023/05/23/microsoft-build-nvidia-ai-windows-rtx/ https://blogs.nvidia.com/blog/2023/05/23/microsoft-build-nvi...
- lelandbatey 3y agoThe comments point out that AMD in the table performing well required the use of Microsoft Olive, and someone in the article comments implies that if you use Microsoft Olive with Nvidia instead of Pytorch with Nvidia, then you'll see the Nvidia jump in performance as well, largely rendering the supposed leap by AMD not relevant. Is that true? Can folks chime in?
- Havoc 3y agoNearly bought one thinking AMD will sort itself out shortly but hard to justify vs a second hand 3090 with 24gb and no cuda hassles
- xigency 3y agoIt makes a lot of sense to invest in a 24GB card for the right price.
- doctorpangloss 3y agoI'm still basking in my good fortune buying a hundred 3090s from crypto miners at rock bottom prices.
- klft 3y ago> Using Microsoft Olive and DirectML instead of the PyTorch pathway results in the AMD 7900 XTX going form a measly 1.87 iterations per second to 18.59 iterations per second! So the headline should be Microsoft Olive vs. PyTorch and not AMD vs. Nvidia.
- mananaysiempre 3y agoThe results of the usual benchmarks are inconclusive between the 7900 XTX and the 4080, Nvidia is only somewhat more expensive, yet CUDA is much more popular than anything AMD is allowed to support. So I’d say this makes sense as an AMD vs Nvidia comparison as well.
- dangus 3y agoThe existence of the 4090 is another issue. I’m not sure which customer willing to spend $1000-1200 to do ML workflows isn’t willing to spend $1600 to get another 20%+ of performance and have the fastest card available. I’m not saying people have unlimited budgets but it just seems like the choice most people in that price range would make.
- selfhoster11 3y agoWhile hobbyist users are not a significant chunk of the market, you can almost bet that they will want the best bang for buck irrespective of the performance of an individual card. Tesla P40 is being heavily discussed in these circles because although it's janky (extremely bad TFLOPS for 16-bit operations), it's still thought to be a decent option for inference due to the high amount of VRAM at its price point.
- 3abiton 3y agoAny references on those P40 discussions?
- Aerroon 3y agoThis raises potential for the next AMD generation though. If they can reach better cost effectiveness but provide more VRAM then that could work.
- brucethemoose2 3y agoWell the problem is that Automatic1111 is not fast... Other diffusers based UIs with PyTorch Triton will net you 40%+ performance. Facebook AITemplate inference in VoltaML will be at least twice as fast as A1111 on a 3080, with support for LORAs, controlnet and such. This supports AMD Instinct cards too. What I am getting at is that people dont really care about A1111 performance on a 3080 because, for the most part, its fast enough.
- kristopolous 3y agoThe extension ecosystem is what makes 1111 the winner for now. SegmentAnything, DreamBooth, ControlNet, OpenPose ... It's almost easy
- brucethemoose2 3y agoSegmentAnything is the big one missing from other UIs, but IMO most of the other extensions are pretty niche, especially with how hackable diffusers is compared to the A1111/Comfy SAI backend.
- Auracle 3y agoI’ve found Adetailer and Regional Prompter to be way better than the Comfy equivalents, sadly.
- cschmid 3y agoCan I also interpret this as: 'AMD's pytorch support is so abysmal that inference is 10x slower than it should be'?
- croes 3y agoShould it not say PyTorch's AMD support?
- dannyw 3y agoIt takes two to tango. AMD is always welcome to contribute patches. You also have to keep in mind some latest gen AMD GPUs don’t even officially support ROCm on Linux. That’s absurd. AMD has a choice to invest more staff into ML support, they’re choosing not to.
- m00x 3y agoFrom Geohotz's investigation in the matter, it doesn't look like it's a manpower issue, it's a quality/culture issue. AMD's firmware GPU division isn't amazing.
- lhl 3y agoI've harped on this on the past, the "official" hardware support list is tiny: https://rocm.docs.amd.com/en/latest/release/gpu_os_support.html#linux-supported-gpus https://rocm.docs.amd.com/en/latest/release/gpu_os_support.h... But, it's worth noting there's different levels of "support." With ROCm 5.6, the 7900XT and 7900XTX RDNA3 cards, while not "officially" supported are gfx1100, which have rocBLAS/MIOpen kernels and work w/o jiggerpokery w/ PyTorch nightly and various HIPified inferencers I've tried like ExLlama.
- DarkmSparks 3y agoBeen watching this quite closely. As far as I summarise, the 7900XTX is the first (and only) desktop GPU from AMD that _might_ be worth buying. (They own the console gaming space, but thats a different story). Not Nvidia beating due to the CUDA issue, but a massive leap in the right direction. Intel is also making _some_ progress with its ARC range. Its going to be happy days for us users if/when AMD/Intel are competitive, and cut some of that monopoly margin off Nvidias pricing, but a way to go yet.
- EtienneK 3y agoThis is not true at all. AMD GPUs have been constantly delivering better bang-for-buck for a while now. Edit: of course I am talking about gaming here since you mentioned consoles
- Iulioh 3y agoEh. For me the problem is technology and not raw performance DLSS and RT are massive for someone with a 4k screen but now 4k gaming hardware outside league of legends lol
- esperent 3y agoDLSS is only marginally better then FSR 2.0. Sure, you can see the difference if you're carefully looking for it in comparison videos. But when you're actually playing it's usually not noticeable to any meaningful degree. You're gonna get occasional artefacts with both. As for ray tracing - to be honest, playing at 4k with an RTX 3070 there's very few games where I can turn it on without the game running unacceptably slow even with DLSS. Ray traced shadows and AO are nice but hardly a deal breaker. I think it's more the desire to have the latest and greatest tech that makes Nvidia cards desirable, rather than any material difference that you'll actually notice while gaming. Some good comparison shots of DLSS 2 and FSR 2 in this video. https://youtu.be/1WM_w7TBbj0 https://youtu.be/1WM_w7TBbj0
- izacus 3y ago
- tamrix 3y agoHow does it compare to nvidias Jetson Orin?
- laserbeam 3y agoI'm actually curious if libraries like pytorch are even trying to move away from CUDA and if moving away from it is worth it. I get why newer ML toolchains would do that but do mainstream established MI frameeorks plan on sticking with nvidia exclusivity for now?
- zwaps 3y agoMove to where? AMD's Rocm doesn't even support their own cards.
- smoldesu 3y agoThere literally aren't other options. Start with Pytorch; if you removed CUDA support, all that's left is hacks honestly. Stuff like Olive and DirectML is equally as proprietary as CUDA while also less-developed, and Metal Performance Shaders are a hacked wrapper for GPU compute. See for yourself: https://pytorch.org/docs/stable/backends.html https://pytorch.org/docs/stable/backends.html There's also ONNX, a Pytorch alternative made by Microsoft that focuses on optimizing for as many targets as possible. It is very difficult to use though, and still under active development. Generally, the issue is that our FAANG companies hate each other too much to do anything about Nvidia. AMD and Apple did have a working relationship (and even a GPGPU acceleration library), but now they are bitter enemies.
- laserbeam 3y agoI'm thinking maybe WebGPU might get one far enough in terms of being able to express ML pipelines? Unsure. I'm not super close to the field.
- delusional 3y agoI've been running pytorch and rocm (5.6 has support for gfx1100 if you compile it yourself) for at least 3 months at 18 it/s on a 7900 XTX. This has been possible for quite a while. Could someone fill me in on what's actually new here, other than the specific technology used?
- Zetobal 3y agoStay away from amd consumer GPUs they are not stable enough... neither hard nor software.
- aranelsurion 3y agoI had the same impression (from personal anecdata) but I wonder if there's any concrete data on it. Do you know of any?
- incrudible 3y agoThis is hard to quantify, there are no benchmarks for how much time people waste to get things to (not) run. In my experience, hardware stability is fine, but software stack has always been problematic. Just the way they treat ROCm makes me stay away. If you have a multi million dollar contract for some datacenter rollout, maybe the opportunity cost amortizes. Maybe.
- lousken 3y agoI am gaming daily on 7900XTX, can you tell me what is not stable? I mean yes, it was not perfect for the first two months I had it, but now I dont have any issues (driver 23.7.1). I have even tried gaming on linux. These days AMD has even better drivers with lower overhead than Nvidia https://youtu.be/H9guEsBly0I https://youtu.be/H9guEsBly0I
- shrx 3y agoWon't using Nvidia cards with Microsoft Olive also provide some boost in performance?