5 ms·
I don't understand. Haven't graphics cards basically been obsolete for deep learning since the first TPUs arrived on the scene in ~2016? Lots of companies are
by peepeepoopoo7 3y ago
I don't understand. Haven't graphics cards basically been obsolete for deep learning since the first TPUs arrived on the scene in ~2016? Lots of companies are offering TPU accelerators now, and it seems like the main thing Nvidia has going for it is momentum. But that doesn't explain this kind of valuation that's hundreds of times greater than their earnings. Personally, it seems a lot like Nvidia is to 2023 what Cisco was to 2000.
- atonse 3y agoI am not an ML expert but as an observer, others have said that why nvidia got right from the beginning, was actually the software support. Stuff like CUDA and good drivers and supporting libraries from over a decade ago? All the libs and researchers and all just use those libs and write software towards it. And as a result it works best on nvidia cards.
- peepeepoopoo7 3y agoBut their valuation is based on forward (future) earnings, using an already obsolete technology.
- tombert 3y agoCorrect me if I'm wrong, but isn't OpenAI still using a ton of Nvidia tech behind the scenes? In addition to GPUs doesn't Nvidia also have dedicated ML hardware?
- jjoonathan 3y agoYes, and all that CUDA software is effectively a moat. ROCm exists but after getting burned badly and repeatedly by OpenCL I'm disinclined to bet on it. At best, my winnings would be avoiding the green tax, at worst, I waste months like I did on OpenCL. That said, AMD used to be in a dire financial situation, whereas now they can afford to fix their shit and actually give chase. NVIDIA has turned the thumb screws very far and they can probably turn them considerably further before researchers jump, but far enough to justify 150x? I have doubts.
- tombert 3y agoThe fact that I hadn't even heard of ROCm until reading your post indicates they got a long way to go to catch up. I've heard of OpenCL but I don't know anyone who actually uses it. I think Apple has something for GPGPU for Metal with performance shaders or compute shaders, but I also don't know of anyone using it for anything, at least not in ML/AI. It's a little irritating that Nvidia has effectively monopolized the GPGPU market so effectively; a part of me wonders if the best that that AMD could do is just make a CUDA-compatibility layer for AMD cards.
- jjoonathan 3y agoIf you look at the ROCm API you'll see that it's pretty much exactly that, a CUDA compatibility layer, but an identical API means little if the functionality behind it has different quirks and caveats and that's harder to assess. I am rooting for ROCm but I can't justify betting on it myself and I suspect most of the industry is in the same boat. For now.
- atonse 3y agoMy question is, is it feasible for AMD to build an ahead of time compiler that transparently translates CUDA instructions into whatever AMD could have, so things Just Work™? They'd be heavily incentivized to do so. Or even put hardware CUDA translators directly into the cards? Or am I misunderstanding CUDA? I think of it as something like OpenGL/DirectX.
- nightski 3y agoIn what universe is it obsolete?
- jjoonathan 3y agoIn the universe of peepeepoopoo7.
- reesul 3y agoAs someone closer to this in the industry (embedded ML.. and trying to compete) I agree with the sentiment. Their software is good, I willingly admit. Porting a model to embedded is hard. With NVIDIA, you basically don’t have to port. This has paid dividends for them, pun not intended. I don’t really see the Nvidia monopoly on ML training stopping anytime soon.
- rfoo 3y agoIt is more expensive (in engineering cost) to port all the world's research (and your own one) to the TPU of your choice, than just paying NVIDIA.
- villgax 3y agoNot one provider apart from GCP has TPUs, they don't have them available for consumers to buy & experiment with. No-one experiments multi-day stuff on the cloud without big pockets or company money especially not PhD students or hobbyists
- peepeepoopoo7 3y agoAWS has their own accelerators, and they're a much better value than their GPU instances.
- popinman322 3y ago1xH100 is faster than 8xA100 for a work-in-progress architecture I'm iterating on. Meanwhile the code doesn't work on TPUs right now because it hangs before initialization. (This is with PyTorch for what it's worth) All that to say Nvidia's hardware work and software evangelization has really paid off-- CUDA just works™ and performance continues to increase. TPUs are good hardware, but TPUs are not available outside of GCP. There's not as much of an incentive for other companies to build software around TPUs like there is with CUDA. The same is likely true of chips like Cerebras' wafer scale accelerators as well. Nvidia's won a stable lead on their competition that's probably not going to disappear for the next 2-5 years and could compound over that time.
- nightski 3y agoIt baffles me to this day that Google never made TPUs more widely available. Then again it is Google...
- ls612 3y agoThey probably saw TPUs as their moat…
- haldujai 3y agoThis has been my experience as well with TPUs and A100s. I haven’t used H100s yet (OOM on 1) but I believe the training throughout benchmarks from Nvidia on transformer workloads is 2.5x from A100s. The effort to make (PyTorch) code run on TPUs is not worth it and my lab would rather rent (discounted) Nvidia GPUs than use free TRC credits we have at the moment. Additionally, at least in Jan 2023 when I last tried this, PyTorch XLA had a significant reduction in throughput so to really take advantage you would probably need to convert to Jax/TF which are used internally at Google and better supported.
- YetAnotherNick 3y agoFor transformer, v4 chip has 70-100% compute capacity and 40% memory of A100 for pretty much the same price. The only benefit is better networking speed for TPU compared to GPU cluster, allowing very large models to scale better, where for GPU model need to fit in NVlink connected GPUs, which is 320 billion parameters for 8*80 GB A100.
- pjc50 3y agoPeople underestimate how meme-y both the stock market and the underlying customer market is. I don't think there's anything like the level of TPUs shipping as there are GPUs? If people end up with an "AI accelerator" in their PC, it would be quite likely to have NVIDIA branding.
- dahart 3y agoThis quote from Wednesday’s TinyCorp article seems apropos: “The current crop of AI chip companies failed. Many of them managed to tape out chips, some of those chips even worked. But not a single one wrote a decent framework to use those chips. They had similar performance/$ to NVIDIA, and way worse software. Of course they failed. Everyone just bought stuff from NVIDIA.” https://geohot.github.io//blog/jekyll/update/2023/05/24/the-tiny-corp-raised-5M.html https://geohot.github.io//blog/jekyll/update/2023/05/24/the-...
- 0xcde4c3db 3y agoFor at least some applications, the details of the processor architecture are dominated by how much high-throughput RAM you can throw at the problem, and GPUs are by far the cheapest and most accessible way of cramming a bunch of high-throughput RAM into a computer. While it's not exactly a mainstream solution, some people have built AI rigs in 2022/2023 with used Vega cards because they're a cheap way to get HBM.
- jandrese 3y agoThe thing that all of these hardware companies don't understand is that it is the software that keeps the boys in the yard. If you don't have something that works as well as CUDA then it doesn't matter how good your hardware is. The only company that seems to understand this is nVidia, and they are the ones eating everyone's lunch. The software side is hard, it is expensive, it takes loads of developer hours and real life years to get right, and it is necessary for success.
- RobotToaster 3y agoI've been wondering the same. With crypto we saw the adoption of ASICs pretty quickly, you would think we would see the same with AI.