16 ms·
An interview with AMD CEO Lisa Su about solving hard problems
- logicchains 2y agoAnd yet they still can't solve the problem of their GPU driver/software stack for ML being much worse than NVidia's. It seems like the first step is easy: pay more for engineers. AMD pays engineers significantly less than NVidia, and it's presumably quite hard to build a competitive software stack while paying so much less. You get what you pay for.
- partypete 2y agoCame here to say this. They only just recently got an AMD GPU on MLPerf thanks to a (different company), Tinycorp by George Hotz. I guess basic ML performance is too hard a problem.
- bee_rider 2y agoI dunno, it a world where hardware companies like, sold hardware, and then software companies wrote software and sold that could be pretty nice. It is cool that Hotz is doing something other than contribute to an anticompetitive company’s moat.
- wordofx 2y agoAnd yet people seem to work just fine with ML on AMD GPUs when they aren’t thinking about Jensen.
- logicchains 2y agoWhich AMD GPUs? Most consumer AMD GPUs don't even support ROCm.
- lhl 2y agoROCm 6.0 and 6.1 list RDNA3 (gfx1100) and RDNA2 (gfx1030) in their supported architectures list: https://rocm.docs.amd.com/en/latest/compatibility/compatibility-matrix.html https://rocm.docs.amd.com/en/latest/compatibility/compatibil... Although "official" / validated support^ is only for PRO W6800/V620 for RDNA2 and RDNA3 RX 7900's for consumer. Based on lots of reports you can probably just HSA_OVERRIDE_GFX_VERSION override for other RDNA2/3 cards and it'll probably just work. I can get GPU-accelerate ROCm for LLM inferencing on my Radeon 780M iGPU for example w/ ROCm 6.0 and HSA_OVERRIDE_GFX_VERSION=11.0.0 (In the past some people also built custom versions of ROCm for older architectures (eg ROC_ENABLE_PRE_VEGA=1) but I have no idea if those work still or not.) ^ https://rocm.docs.amd.com/projects/install-on-linux/en/latest/reference/system-requirements.html https://rocm.docs.amd.com/projects/install-on-linux/en/lates...
- JonChesterfield 2y agoDebian, Arch and Gentoo have ROCm built for consumer GPUs. Thus so do their derivatives. Anything gfx9 or later is likely to be fine and gfx8 has a decent chance of working. The https://github.com/ROCm/ROCm https://github.com/ROCm/ROCm source has build scripts these days. At least some of the internal developers largely work on consumer hardware. It's not as solid as the enterprise gear but it's also very cheap so overall that seems reasonable to me. I'm using a pair of 6900XT, with a pair of VII's in a backup machine. For turn key proprietary stuff where you really like the happy path foreseen by your vendor, in classic mainframe style, team green is who you want.
- paulmd 2y ago> For turn key proprietary stuff where you really like the happy path foreseen by your vendor there really was no way for AMD to foresee that people might want to run GPGPU workloads on their polaris cards? isn't that a little counterfactual to the whole OpenCL and HSA Framework push predating that? Example: it's not that things like Bolt didn't exist to try and compete with Thrust... it's that the NVIDIA one has had three updates in the last month and Bolt was last updated 10 years ago. You're literally reframing "having working runtime and framework support for your hardware" as being some proprietary turnkey luxury for users, as well as an unforeseeable eventuality for AMD. It wasn't a development priority, but users do like to actually build code that works etc. That's why you got kicked to the curb by Blender - your OpenCL wasn't stable even after years of work from them and you. That's why you got kicked to the curb by Octane - your Vulkan Compute support wasn't stable enough to even compile their code successfully. That's the story that's related by richg42 about your OpenGL driver implementation too - that it's just paper features and resume-driven development by developers 10 years departed all the way down. The issues discussed by geohotz aren't new, and they aren't limited to ROCm or deep learning in general. This is, broadly speaking, the same level of quality that AMD has applied to all its software for decades. And the social-media "red team" loyalism strategy doesn't really work here, you can't push this into "AMD drivers have been good for like 10 years now!!!" fervor when the understanding of the problems are that broad and that collectively shared and understood. Every GPGPU developer who's tried has bounced off this AMD experience for literally an entire generation running now. The shared collective experience is that AMD is not serious in the field, and it's difficult to believe it's a good-faith change and interest in advancing the field rather than just a cashgrab. It's also completely foreseeable that users want broad, official support for all their architectures, and not one or two specifics etc. Like these aren't mysteries that AMD just accidentally forgot about, etc. They're basic asks that you are framing as "turnkey proprietary stuff", like a working opencl runtime or a working spir-v compiler. What was it linus said about the experience of working with NVIDIA? That's been the experience of the GPGPU community working with AMD, for decades. Shit is broken and doesn't work, and there's no interest in making it otherwise. And the only thing that changed it is a cashgrab, and a working compiler/runtime is still "turnkey proprietary stuff" they have to be arm-twisted into doing by Literally Being Put On Blast By Geohotz Until It's Fixed. "Fuck you, AMD" is a sentiment that there is very valid reasons to feel given the amount of needless suffering you have generated - but we just don't do that to red team, do we? But you guys have been more intransigent about just supporting GPGPU, no matter what framework, please just pick one, get serious and start working already than NVIDIA ever was about wayland etc. You've blown decades just refusing to ever shit or get off the pot (without even giving enough documentation for the community to just do it themselves). And that's not an exaggeration - I bounced off the AMD stack in 2012, and it wasn't a new problem then either. It's too late for "we didn't know people wanted a working runtime or to develop on gaming cards" to work as an excuse, after decades of overt willing neglect it's just patronizing. Again, sorry, this is ranty, it's not that I'm upset at you personally etc, but like, my advice as a corporate posture here is don't go looking for a ticker-tape parade for finally delivering a working runtime that you've literally been advertising support for for more than a decade like it's some favor to the community. These aren't "proprietary turnkey features" they're literally the basics of the specs you're advertising compliance with, and it's not even just one it's like 4+ different APIs that have this problem with you guys that has been widely known, discussed in tech blogs etc for more than a decade (richg42). I've been saying it for a long time, so has everyone else who's ever interacted with AMD hardware in the GPGPU space. Nobody there cared until it was a cashgrab, actually half the time you get the AMD fan there to tell you the drivers have been good for a decade now (AMD cannot fail, only be failed). It's frustrating. You've poisoned the well with generations of developers, with decades of corporate obstinance that would make NVIDIA blush, please at least have a little contrition about the whole experience and the feelings on the other side here.
- FeepingCreature 2y agoI have a 7900 XTX. There's a known firmware crash issue with ComfyUI. It's been reported like a year ago. Every rocm patch release I check the notes, and every release it goes unfixed. That's not to go into the intense jank that is the rocm debian repo. If we need DL at work, I'll recommend Nvidia, no question.
- sitkack 2y agoEveryone does software poorly, hardware companies more so.
- jajko 2y agoWell this is glaringly obvious to whole world, and Nvidia managed to get it right. Surely a feat that can be repeated elsewhere when enough will is spread over some time. And it would make them grow massively, something no shareholder ever frowns upon.
- rf15 2y ago> Nvidia managed to get it right I don't think they did. If you work in the space and watched it develop over the years, you can see that there's been (and still are) plenty of jank and pain points to be found. Their spread mostly comes from their dominant market position before, afaik.
- modeless 2y agoThey succeeded not because they are perfect but because they are the least bad. By far.
- DanielHB 2y agoAlso CUDA first released in 2007, I remember researching GPUs at the time and wondering what the hell CUDA was (I was a teenager). They were VERY early to the party and has A LOT of time to improve. Everyone is catching up to ~10 years of a headstart.
- croes 2y agoOr people are just used to the same kind of bad
- cepth 2y agoA couple of thoughts here. * AMD's traditional target market for its GPUs has been HPC as opposed to deep learning/"AI" customers. For example, look at the supercomputers at the national labs. AMD has won quite a few high profile bids with the national labs in recent years: - Frontier (deployment begun in 2021) (https://en.wikipedia.org/wiki/Frontier_(supercomputer) https://en.wikipedia.org/wiki/Frontier_(supercomputer)) - used at Oak Ridge for modeling nuclear reactors, materials science, biology, etc. - El Capitan (2023) (https://en.wikipedia.org/wiki/El_Capitan_(supercomputer) https://en.wikipedia.org/wiki/El_Capitan_(supercomputer)) - Livermore national lab AMD GPUs are pretty well represented on the TOP500 list (https://top500.org/lists/top500/list/2024/06/ https://top500.org/lists/top500/list/2024/06/), which tends to feature computers used by major national-level labs for scientific research. AMD CPUs are even moreso represented. * HPC tends to focus exclusively on FP64 computation, since rounding errors in that kind of use-case are a much bigger deal than in DL (see for example https://hal.science/hal-02486753/document https://hal.science/hal-02486753/document). NVIDIA innovations like TensorFloat, mixed precision, custom silicon (e.g., the "transformer engine") are of limited interest to HPC customers. It's no surprise that AMD didn't pursue similar R&D, given who they were selling GPUs to. * People tend to forget that less than a decade ago, AMD as a company had a few quarters of cash left before the company would've been bankrupt. When Lisa Su took over as CEO in 2014, AMD market share for all CPUs was 23.4% (even lower in the more lucrative datacenter market). This would bottom out at 17.8% in 2016 (https://www.trefis.com/data/companies/AMD,.INTC/no-login-required/TfU13vlN/Intel-vs-AMD-How-Have-Revenues-Key-Operating-Metrics-Changed-Over-Recent-Years https://www.trefis.com/data/companies/AMD,.INTC/no-login-req...). AMD's "Zen moment" didn't arrive until March 2017. And it wasn't until Zen 2 (July 2019), that major datacenter customers began to adopt AMD CPUs again. * In interviews with key AMD figures like Mark Papermaster and Forrest Norrod, they've mentioned how in the years leading up to the Zen release, all other R&D was slashed to the bone. You can see (https://www.statista.com/statistics/267873/amds-expenditure-on-research-and-development-since-2001 https://www.statista.com/statistics/267873/amds-expenditure-...) that AMD R&D spending didn't surpass its previous peak (on a nominal dollar, not even inflation-adjusted, basis) until 2020. There was barely enough money to fund the CPUs that would stop the company from going bankrupt, much less fund GPU hardware and software development. * By the time AMD could afford to spend on GPU development, CUDA was the entrenched leader. CUDA was first released in 2003(!), ROCm not until 2016. AMD is playing from behind, and had to make various concessions. The ROCm API is designed around CUDA API verbs/nouns. AMD funded ZLUDA, intended to be a "translation layer" so that CUDA programs can run as a drop-in on ROCm. * There's a chicken-and-egg problem here. 1) There's only one major cloud (Azure) that has ready access to AMD's datacenter-grade GPUs (the Instinct series). 2) I suspect a substantial portion of their datacenter revenue still comes from traditional HPC customers, who have no need for the ROCm stack. 3) The lack of a ROCm developer ecosystem means that development and bug fixes come much slower than they would for CUDA. For example, the mainline TensorFlow release was broken on ROCm for a while (you had to install the nightly release). 4) But, things are improving (slowly). ROCm 6 works substantially better than ROCm 5 did for me. PyTorch and TensorFlow benchmark suites will run. Trust me, I share the frustration around the semi-broken state that ROCm is in for deep learning applications. As an owner of various NVIDIA GPUs (from consumer laptop/desktop cards to datacenter accelerators), in 90% of cases things just work on CUDA. On ROCm, as of today it definitely doesn't "just work". I put together a guide for Framework laptop owners to get ROCm working on the AMD GPU that ships as an optional add-in (https://community.frame.work/t/installing-rocm-hiplib-on-ubuntu-22-04/50412/2 https://community.frame.work/t/installing-rocm-hiplib-on-ubu...). This took a lot of head banging, and the parsing of obscure blogs and Github issues. TL;DR, if you consider where AMD GPUs were just a few years ago, things are much better now. But, it still takes too much effort for the average developer to get started on ROCm today.
- Bancakes 2y agoExactly. ROCm is only available for the top tier RX 7900 GPUs and you’re expected to run Linux. AMD was (is) tinkering with a translation layer for CUDA, much like how WINE translates directX. Great idea but it’s been taking a while in this fast paced market.
- drbscl 2y ago> AMD was (is) tinkering with a translation layer for CUDA From what I understand, they dropped the contract with the engineer who was working on it. Fortunately, as part of the contract, said engineer stipulated that the project would become open source, so now it is, and is still being maintained by that engineer.
- anticensor 2y ago> ROCm is only available for the top tier RX 7900 GPUs and you’re expected to run Linux. Fixed it for you: ROCm is only officially supported for the top tier RX 7900 GPUs and you’re expected to run Linux. Desktop class cards work if you apply a "HSA version override".
- Bancakes 2y agoCool. I was thinking of getting a 7800 XT over a 4070. Hope I can get llama 70B working nearly as well.
- refurb 2y agoI’m not anti-CEO, I think they play an important role, but why would you interview a CEO about hard technological problems?
- yen223 2y agoLisa Su is one of the (surprisingly rare) tech CEOs who comes from an engineering background.
- automatic6131 2y agoAnd so is her first cousin (once removed) Jensen Huang of Nvidia!
- drexlspivey 2y agoYou mean like the CEOs of Intel, Nvidia, TSMC, ASML, Micron etc?
- user_7832 2y agoIn general it's very common for CEOs to be of a business/MBA background. Apple and Google both have this, along with tons of other companies outside the semiconductor industry.
- scarface_74 2y agoApple as a hardware company needed someone who was good at operations and logistics. Tim Cook is one of the best in the business. On the other hand, I have no idea what Google’s CEOs purpose is. Google has no vision and can’t produce new products to save their lives. Google is unique among the current BigTech companies in being completely incapable of diversifying or pivoting to new market opportunities
- MrBuddyCasino 2y agoGoogle might be the worst run unicorn. How they fell so low from that high is truly an achievement. They will continue to print money, but to think Google was once programmer Utopia and Googlers were regarded as demi-goods... inconceivable now.
- djvu97 2y agoWas I the only one who was expecting Lisa Su to attempt Leetcode hard in this interview?
- blitzar 2y agoIt is an interview with a CEO - expect platitudes and fortune cookie wisdom that contradict their own (corporate) behaviours.
- stuxnet79 2y agoAs an IC this is exactly what makes 'fireside chats' with executives so unsatisfying.
- pizzalife 2y agoThere's also never an actual fire. Like at least put a video of a fireplace on in the background.
- lupire 2y agoIt's called "fireside chat" because of the layoffs. :-/
- baobabKoodaa 2y agoFor the people downvoting this joke: the joke is that "fireside" means "the side of the people who are doing the firing". I think it's funny.
- altdataseller 2y agoNot really. I wouldn't classify leetcode as "hard problems". Maybe dumb problems that don't really help anyone, but no, not "hard problems"
- scottyah 2y ago
- NKosmatos 2y agoI was todays years old, when I found out that Lisa Su (CEO of AMD) and Jensen Huang (co-founder and CEO of NVIDIA) are relatives! If you can't do a merge, it's good to have family onboard ;-)
- adverbly 2y agoI couldn't believe this so I had to Google it. First cousins once removed. I don't really know what to think about this...
- jasonvorhe 2y agoAll US presidents are directly related to each other: https://curiousmindmagazine.com/all-us-presidents-including-trump-are-descendants-of-the-same-english-king/ https://curiousmindmagazine.com/all-us-presidents-including-... The world's a stage. :)
- tsunamifury 2y ago… and you ain’t on it.
- SketchySeaBeast 2y agoA relative from 800 years ago doesn't honestly seem that impressive. That's ~30 generations? That's a whole lot of people. Good luck finding a venue large enough for that family reunion.
- Retric 2y agoIt’s more meaningful than you might think because the majority of people on earth probably aren’t decedents of this guy. Especially in the 1700’s when our first presidents where born. Essentially zero US presidents are Asian, Hispanic, etc. Pick a Native American from 30 generations ago and you don’t see this kind of family tree. It’s an expanded circle of privilege through time.
- arnon 2y ago> please don’t think of AMD as an x86 company, we are a computing company, and we will use the right compute engine for the right workload. Love that quote.
- high_na_euv 2y agoCrazy that they need to write such things because incompetent investors think that ISA has huge implications
- leonheld 2y agoThe shocker is: a lot of engineers think that the ISA has huge implications.
- chasil 2y agoThere was an interview with one of the SPARC creators who said that a huge benefit of control of the instruction set was the ability to take the platform in directions that (Intel) resellers could not. He was otherwise largely agnostic on the benefits of SPARC.
- paulmd 2y agoISA shouldn’t be more than a 10% or 20% performance impact or so, per jim Keller
- SkyMarshal 2y agoThat's not trivial...
- nextaccountic 2y agoWhat this means is that a better ISA might permanently be one generation ahead in terms of performance
- glitchc 2y agoISAs do have huge implications. A poorly designed ISA can balloon the instruction count, which has a direct impact on program size and the call stack.
- sigmoid10 2y agoOmg. I know this is mostly marketing speaking, but this is her reply when asked about AMD's reticence to software: > Well, let me be clear, there’s no reticence at all. [...] I think we’ve always believed in the importance of the hardware-software linkage and really, the key thing about software is, we’re supposed to make it easy for customers to use all of the incredible capability that we’re putting in these chips, there is complete clarity on that. I'm baffled how clueless these CEOs sometimes seem about their own product. Like, do you even realize that this the reason why Nvidia is mopping the floor with your stuff? Have you ever talked to a developer who had to work with your drivers and stack? If you don't start massively investing on that side, Nvidia will keep dominating despite their outrageous pricing. I really want AMD to succeed here, but with management like that I'm not surprised that they can't keep up. Props to the interviewer for not letting her off the hook on this one after she almost dodged it.
- cherryteastain 2y agoWhat is she supposed to say? Perhaps "our products have bad software, don't buy them, go buy Nvidia instead"?
- sigmoid10 2y agoShe could admit that they fell behind on this one and really need to focus on closing the gap now. But instead she says it's all business as usual, which assures me that I won't give their hardware another shot for quite a while.
- deleted 2y ago[deleted]
- wslh 2y agoIt is good to mention the AMD's Steam Deck CPU [1] running in the Steam Deck [2] and not less important that the Steam Deck also has Linux (and KDE) incorporated [3]. [1] https://www.techpowerup.com/cpu-specs/steam-deck-cpu-lcd.c3397 https://www.techpowerup.com/cpu-specs/steam-deck-cpu-lcd.c33... [2] https://store.steampowered.com/steamdeck https://store.steampowered.com/steamdeck [3] https://help.steampowered.com/en/faqs/view/671A-4453-E8D2-323C https://help.steampowered.com/en/faqs/view/671A-4453-E8D2-32...
- rhelz 2y agoI have endless admiration for Lisa Su, but lets be honest, the reason AMD and Nvidia are so big today is that Intel has had amazingly bad management since about 2003. They massively botched the 32-bit to 64-bit transition....I did some work for the first Itanium, and everybody knew even before it came out that it would never show a profit. And they were contractually obligated to HP to not make 64-bit versions of the x86....so we just had to sit there and watch while AMD beat us to 1 Gigahertz, and had the 64-bit x86 market to itself.... When they fired Pat Gelsinger, their doom was sealed. Thank God they hired him back, but now they are in the same position AMD and Nvidia used to be in: Intel just has to wait for Nvidia and AMD to have bad management for two straight decades....
- K0SM0S 2y agoI don't mean to take away from Intel's underwhelming management. But regardless, Keller's Athlon 64 or Zen are great competitors. Likewise, CUDA is Nvidia's massive achievement. The growth strategy of that product (involving lots of free engineer hours given to clients on-site) deserves credit.
- rhelz 2y ago// I don't mean to take away from Intel's underwhelming management chuckle lets give full credit where credit is due :-) Athlon was an epochal chip. Here's the thing though---if you are a market leader, one who was as dominant as Intel was, it doesn't matter what the competition does, you have the power to keep dominating them by doing something even more epochal. That's why it can be so frustrating working for a #2 or #3 company....you are still expected to deliver epochal results like clockwork. But even if you do, your success is completely out of your hands. Bringing out epochal products doesn't get you ahead, it just lets you stay in the game. Kind of like the Red Queen in Alice in Wonderland. You have to run as fast as you can just to stay still. All you can do is try to stay in the game long enough until the #1 company makes a mistake. If #1 is dominate enough, they can make all kinds of mistakes and still stay on top, just by sheer market inertia. Intel was so dominate that it took DECADES of back-to-back mistakes to lose its dominate position. Intel flubbed the 32-64 bit transition. On the low end, it flubbed the desktop to mobile transition. On the high end, it flubbed the CPU-GPU transition. Intel could have kept its dominate position if it had only flubbed one of them. But from 2002 to 2022, Intel flubbed every single transition in the market. Its a measure of just how awesome Intel used to be that it took 20 years....but there's only so many of those that you can do back-to-back and still stay #1.
- pbrum 2y agoThis article inspired me to check out the respective market caps for Intel and AMD. Hell of a turnaround! I remember the raging wars between the Pentiums and the Athlons. Intel won. The GPU wars between Nvidia and ATI; Nvidia won. Thereafter AMD, the supposed loser to Intel, absorbed ATI, the supposed loser to Nvidia. But I love that the story didn't end there. Look at what AMD did with Sony PlayStation (extensively discussed in this interview)...and that's without getting into the contemporary GPU AI revolution that's likely driving AMD's $250+ billion market cap. Epic!