4 ms·
Do we expect special AI processors to diverge from GPUs? Like, processors that can do parallel neural network computations but cannot draw graphics?
by Buttons840 7mo ago
Do we expect special AI processors to diverge from GPUs? Like, processors that can do parallel neural network computations but cannot draw graphics?
- dagmx 7mo agoThat’s already the norm no? Pretty much every hardware vendor has an NPU
- c0balt 7mo agoThat is already the case with datacenter "GPUs". A A100, MI300 or Intel PVC/Gaudi does not have useful graphics performance nor capabilities. Coprocessors ala NPU/VPU are also on the rise again for CPUs.
- elcritch 7mo agoGreat now I’m envisioning a rich guy using an A100 as his desktop GPU just to show off. Which begs the question if that’s even possible.
- userbinator 7mo agoIt has no video output.
- coffeebeqn 7mo agoI believe some cards at least you can make the motherboard display ports the output
- vel0city 7mo agoThis is kind of true back in the day though. Uninformed people would buy Quadro cards because they were the most expensive GPU on Newegg only to realize this thing sucks for gaming.
- jiggawatts 7mo agoYes. Even the latest NVIDIA Blackwell GPUs are general purpose, albeit with negligible "graphics" capabilites. They can run fairly arbitrary C/C++ code with only some limitations, and the area of the chip dedicated to matrix products (the "tensor units") is relatively small: less than 20% of the area! Conversely, the Google TPUs dedicate a large area of each chip to pure tensor ops, hence the name. This is partly why Google's Gemini is 4x cheaper than OpenAI's GPT5 models to serve. Jensen Huang has said in recent interviews that he stands by the decision to keep the NVIDIA GPUs more general purpose, because this makes them flexible and able to be adapted to future AI designs, not just the current architectures. That may or may not pan out. I strongly suspect that the winning chip architecture will have about 80% of its area dedicated to tensor units, very little onboard cache, and model weights streamed in from High Bandwidth Flash (HBF). This would be dramatically lower power and cost compared to the current hardware that's typically used. Something to consider is that as the size of matrices scales up in a model, the compute needed to perform matrix multiplications goes up as the cube of their size, but the other miscellaneous operations such as softmax, relu, etc.. scale up linearly with the size of the vectors being multiplied. Hence, as models scale into the trillions of parameters, the matrix multiplications ("tensor" ops) dominate everything else.
- IsTom 7mo agoI'm not following the whole LLM space, but > the compute needed to perform matrix multiplications goes up as the cube of their size, are they really not using even Strassen multiplication?
- jiggawatts 7mo agoAFAIK the best practical matrix multiplication algorithms scale as roughly N^2.7 which is close enough to N^3 to not matter for the point that I'm trying to make.
- jcranmer 7mo agoI'm not aware of any major BLAS library that uses Strassen's algorithm. There's a few reasons for this; one of the big ones is Strassen is much worse numerical performance than traditional matrix multiplication. Another big one is that at very large dense matrices--which are using various flavors of parallel algorithms--Strassen vastly increases the communication overhead. Not to mention that the largest matrices are probably using sparse matrix arithmetic anyways, which is a whole different set of algorithms.
- pjmlp 7mo agoYes, this has already been the case for years on mobile devices, CoPilot+ PC design requires this approach as well. Additionally, GPUs are going back to the early days, by becoming general purpose parallel compute devices, where you can use the old software rendering techniques, now hardware accelerated.
- PunchyHamster 7mo agoWe kinda already have it with NPU/TPUs, tho they are usually attached to CPUs (and for some reason come with near zero proper documentation). I can see separate cards for datacenter use but for consumers they will probably come on same SOC as CPU
- amelius 7mo agoI'm expecting in the not so distant future we'll even have LLMs baked into an ASIC, in our PCs.