6 ms·
Bill Dally from Nvidia argues that there is "no gain in building a specialized accelerator", in part because current overhead on top of the arithmetic is in the
by celrod 2y ago
Bill Dally from Nvidia argues that there is "no gain in building a specialized accelerator", in part because current overhead on top of the arithmetic is in the ballpark of 20% (16% of IMMA and 22% for HMMA units)
https://www.youtube.com/watch?v=gofI47kfD28 https://www.youtube.com/watch?v=gofI47kfD28
- AnthonyMouse 2y agoThere does seem to be a somewhat obvious advantage: If all it has to do is matrix multiplication and not every other thing a general purpose GPU has to be good at then it costs less to design. So now someone other than Nvidia or AMD can do it, and then very easily distinguish themselves by just sticking a ton of VRAM on it. Which is currently reserved for GPUs that are extraordinarily expensive, even though the extra VRAM doesn't cost a fraction of the price difference between those and an ordinary consumer GPU.
- bjornsing 2y agoExactly. And that means you not only save the 22% but also a large chunk of the Nvidia margin.
- Animats 2y agoAnd, sure enough, there's a new AI chip from Intellifusion in China that's supposed to be 90% cheaper. 48 TOPS in int8 training performance for US$140.[1] [1] https://www.tomshardware.com/tech-industry/artificial-intelligence/chinese-chipmaker-launches-14nm-ai-processor-thats-90-cheaper-than-gpus https://www.tomshardware.com/tech-industry/artificial-intell...
- pfdietz 2y agoI wonder what the cost of power to run these chips is. If the power cost ends up being large compared to the hardware cost, it could make sense to buy more chips and run them when power is cheap. They could become a large source of dispatchable demand.
- pclmulqdq 2y agoInt8 training has very few applications, and int8 ops generally are very easy to implement. Int8 is a decent inference format, but supposedly doesn't work well for LLMs that need a wide dynamic range.
- papruapap 2y agoI really hope we see AI-PU (or with some other name, INT16PU, why not) for the consumer market sometime soon. Or been able to expand GPU memory using a pcie socket (not sure if technically possible).
- hhsectech 2y agoIsn't this what resizeable BAR and direct storage are for?
- PeterisP 2y agoThe while point of GPU memory is that it's faster to access than going to memory (like your main RAM) through the PCIe bottleneck.
- deleted 2y ago[deleted]
- throwaway4aday 2y agoMy uninformed question about this is why can't we make the VRAM on GPUs expandable? I know that you need to avoid having the data traverse some kind of bus that trades overhead for wide compatibility like PCIe but if you only want to use it for more RAM then can't you just add more sockets whose traces go directly to where they're needed? Even if it's only compatible with a specific type of chip it would seem worthwhile for the customer to buy a base GPU and add on however much VRAM they need. I've heard of people replacing existing RAM chips on their GPUs[0] so why can't this be built in as a socket like motherboards use for RAM and CPUs? [0] https://www.tomshardware.com/news/16gb-rtx-3070-mod https://www.tomshardware.com/news/16gb-rtx-3070-mod
- carbotaniuman 2y agoReplacing RAM chips on GPUs involves resoldering and similar things - those (for the most part) maintain the signal integrity and performance characteristics of the original RAM. Adding sockets complicates the signal path (iirc), so it's harder for the traces to go where they're needed, and realistically given a trade-off between speed/bandwidth and expandability I think the market goes with the former.
- WithinReason 2y agoDesigning it is easy and always has been. Programming it is the bottleneck. Otherwise Nvidia wouldn't be in the lead.
- markhahn 2y agobut programming it is "import pytorch" - nothing nvidia-specific there. the mass press is very impressed by Cuda, but at least if we're talking AI (and this article is, exclusively), it's not the right interface. and in fact, Nv's lead, if it exists, is because they pushed tensor hardware earlier.
- achierius 2y agoSomeone does, in fact, have to implement everything underneath that `import` call, and that work is _very_ hard to do for things that don't closely match Nvidia's SIMT architecture. There's a reason people don't like using dataflow architectures, even though from a pure hardware PoV they're very powerful -- you can't map CUDA's, or Pytorch's, or Tensorflow's model of the world onto them.
- WithinReason 2y agoI'm talking about adding Pytorch support for your special hardware. Nv's lead is due to them having Pytorch support.
- KaoruAoiShiho 2y agoEh if you're running in production you'll want something lower level and faster than pytorch.
- cma 2y agoThere are other operations for things like normalization in training, which is why most successful custom stuff has focused on inference I think. As architectures changed and needed various different things some custom built training hardware got obsoleted, Keller talked about that affecting Tesla's Dojo and making it less viable (they bought a huge nvidia cluster after it was up). I don't know if TPU ran into this, or they made enough iterations fast enough to keep adding what they needed as they needed it.
- ericye16 2y agoAI models are not all matrix multiplications, and they tend to involve other operations. Also, they change super fast, much faster than hardware cycles, so if your hardware isn't general-purpose enough, the field will move past you and obsolete your hardware before it comes out.
- AnthonyMouse 2y agoAI models are mostly matrix multiplications and have been that way for a few years now, which is longer than a hardware cycle. Moreover, if the structure changes then the hardware changes regardless of whether it's general purpose or not, because then it has to be optimized for the new structure. Everybody cares about VRAM right now yet you can get a P40 with 24GB for 10% of the price of a 24GB RTX 4090. Why? No tensor cores, the things used for matrix multiplication.