7 ms·
128GB in one chip seems important with the rise of sparse architectures like MoE. Hopefully these are competitive with Nvidia's offerings, though in the end the
by rileyphone 2y ago
128GB in one chip seems important with the rise of sparse architectures like MoE. Hopefully these are competitive with Nvidia's offerings, though in the end they will be competing for the same fab space as Nvidia if I'm not mistaken.
- latchkey 2y agoAMD MI300x is 192GB.
- tucnak 2y agoWhich would be impressive had it _actually_ worked for ML workloads.
- Hugsun 2y agoDoes it not work for them? Where can I learn why?
- tucnak 2y agoJust go have a look around Github issues in their ROCm repositories on Github. A few months back the top excuse re: AMD was that we're not supposed to use their "consumer" cards, however the datacenter stuff is kosher. Well, guess what, we have purchased their datacenter card, MI50, and it's similarly screwed. Too many bugs in the kernel, kernel crashes, hangs, and the ROCm code is buggy / incomplete. When it works, it works for a short period of time, and yes HBM memory is kind of nice, but the whole thing is not worth it. Some say MI210 and MI300 are better, but it's just wishful thinking as all the bugs are in the software, kernel driver, and firmware. I have spent too many hours troubleshooting entry-level datacenter-grade Instinct cards with no recourse from AMD whatsoever to pay 10+ thousands for MI210 a couple-year old underpowered hardware, and MI300 is just unavailable. Not even from cloud providers which should be telling enough.
- latchkey 2y ago[flagged]
- buildbot 2y agoProbably not, I have had the same experience with a 780m and a Mi60...
- Scaevolus 2y agoEffectively everyone that has attempted to use AMD hardware for ML comes away with these opinions, the main difference is how angrily they express it.
- sangnoir 2y agoSo who's buying all the MI300s? Groq seems to be fine with AMD.
- tucnak 2y agoGroq doesn't use AMD afaik, they had designed hardware of their own, which is actually 1000s of SRAM chips in a trench-coat.
- latchkey 2y agoSpeaking of SRAM, I found this relevant comment insightful: https://news.ycombinator.com/item?id=39966620 https://news.ycombinator.com/item?id=39966620
- latchkey 2y ago> So who's buying all the MI300s? We are. Just closed additional funding and getting quotes now.
- latchkey 2y agoWe are the 4th non-hyperscaler business on the planet to even get access to MI300x and we just got it in early March. From what I understand, hyperscalers have had fantastic uptake of this hardware. I find it hard to believe "everyone" comes away with these opinions. https://www.evp.cloud/post/diving-deeper-insights-from-our-llm-inference-testing https://www.evp.cloud/post/diving-deeper-insights-from-our-l...
- huac 2y agoThere's a number of scaled AMD deployments, including Lamini (https://www.lamini.ai/blog/lamini-amd-paving-the-road-to-gpu-rich-enterprise-llms https://www.lamini.ai/blog/lamini-amd-paving-the-road-to-gpu...) specifically for LLM's. There's also a number of HPC configurations, including the world's largest publicly disclosed supercomputer (Frontier) and Europe's largest supercomputer (LUMI) running on MI250x. Multiple teams have trained models on those HPC setups too. Do you have any more evidence as to why these categorically don't work?
- latchkey 2y ago> Do you have any more evidence as to why these categorically don't work? They don't. Loud voices parroting George, with nothing to back it up. Here are another couple good links: https://www.evp.cloud/post/diving-deeper-insights-from-our-llm-inference-testing https://www.evp.cloud/post/diving-deeper-insights-from-our-l... https://www.databricks.com/blog/training-llms-scale-amd-mi250-gpus https://www.databricks.com/blog/training-llms-scale-amd-mi25...