9 ms·
Chiplet ASIC supercomputers for LLMs like GPT-4
- steve_mcdougall 3y ago[dead]
- SCUSKU 3y ago94x cost improvement over GPU and 15x TPU is insane, but fits right in there with performance gains seen in Moore's Law. This development presents a more compelling case that we are in fact on the precipice of larger LLMs being able to serve everyone for cheap. Still not really convinced by the AGI argument, but this does spook me. Overall though very cool.
- jacquesm 3y agoIt's insane because it is theoretical. They haven't shown that it works, think of this paper as a prelude to a funding round or research grant so they have to show some kind of advantage. Which I'm highly skeptical of, usually when papers show this kind of improvement over SOTA it tends to be either a mistake or purposeful nonsense.
- kraken12 3y agoYeah, a preliminary architectural study to sanity check if an idea could potentially pay off.
- mathisfun123 3y ago>prelude to a funding round or research grant Group at my school recently got a grant for 10MM for such a fantasy. All they had was an ISA - no RTL, no functional model, no compiler. Kid in my group (co-advised) is busy scrawling assembly on notebook paper lol. Suffice it to say I don't have high hopes for a tapeout anytime soon.
- msoad 3y agoThey claim the investment will be justified for a 1.5 year life span of the system. But LLMs are changing and improving at a much faster speed that 1.5 years feels like centuries!
- thelittleone 3y ago"Moving fast" may take on a whole new meaning and I'd put money on the rate of iteration soon being beyond the vast majority's comprehension (myself included).
- 128bytes 3y agoit's already beyond my comprehension, i've not lived very long but in the time i have i've never seen any technology develop so rapidly and at such a rapidly increasing pace. I assume this is what it must have felt like during the dawn of the age of computing.
- jacquesm 3y agoIt makes you wonder if those singularity proponents don't have a point, and it all depends on whether it keeps accelerating or whether it will slow down again. I hope for the latter and I fear for the former. Even if it does slow down eventually a long enough period of such change is going to make the industrial revolution (whose negative effects we are still coming to terms with today!) like a walk in the park.
- ChatGTP 3y agoUnloading Ray Kurzweil to the cloud in 5, 4, 3…
- keenmaster 3y agoReality is biased towards the fast-moving scenario, so long as we aren’t running into the bounds of physics, which as far as I can tell we’re not. Kurzweil was much more right than he was wrong. The opposite is true of people who strongly disagreed with him and called him a quack.
- seydor 3y agonice, soon i will have my Pocket best friend
- alain94040 3y agoThe key point: A key architectural feature to achieve this is the ability to fit all model parameters inside the on-chip SRAMs of the chiplets to eliminate bandwidth limitations. Doing so is non-trivial as the amount of memory required is very large and growing for modern LLMs ... On-chip memories such as SRAM have better read latency and read/write energy than external memories such as DDR or HBM but require more silicon per bit. We show this design choice wins in the competition of TCO per performance for serving large generative language models but requires careful consideration with respect to the chiplet die size, chiplet memory capacity and total number of chiplets to balance the fabrication cost and model performance (Sec.3.2.2) We observe that the inter-chiplet communication issues can be effectively mitigated through proper software-hardware co- design leveraging mapping strategies such as tensor and pipeline model parallelism
- cubefox 3y agoLarge language models would need tens or hundreds of gigabytes of SRAM. Pretty sure the enormous cost for this makes the approach economically unfeasible.
- senttoschool 3y agoCerebras Wafer Scale Engine has 40GB of onboard SRAM using TSMC 7nm. It uses the entire wafer as the chip. Costs millions per chip though. Source: https://www.anandtech.com/show/16626/cerebras-unveils-wafer-scale-engine-two-wse2-26-trillion-transistors-100-yield https://www.anandtech.com/show/16626/cerebras-unveils-wafer-...
- skummetmaelk 3y agoWafer scale integration is a dead end technology. The engineering issues are just too great.
- senttoschool 3y agoDo you have a source?
- why_only_15 3y agoThis seems like a pretty bad paper. Their headline claim that they are 300x faster than an A100 at serving GPT-3 uses obviously wrong numbers for how fast A100s can run GPT-3. They seem to have misread the DeepSpeed Inference paper and claim that the best throughput on GPT-3 sized models was 18 tok/s, but if you look at figure 8 on page 11 of the paper [1], it shows that they are able to achieve ~74 teraflops on serving LM-175B (GPT-3), which is about 211 tok/s. To calculate TCO, we can say 211 tok/s = 760,000 tok/hour and A100s are about $1/hr, so TCO per 1k tokens using that paper's method is about $0.0013, much lower than the $0.02 that they claim, reducing the claimed TCO advantage from 94x to 6.2x. Combine this with the fact that the paper they used is from a year ago and there are more efficient inference methods than there were then and the speedup probably goes even lower, maybe to 3x. This is without even looking at the chip design itself, whose costs are probably far underestimated. [1]: https://arxiv.org/pdf/2207.00032.pdf https://arxiv.org/pdf/2207.00032.pdf
- jacquesm 3y agoThe whole thing is imaginary: "In this paper, we propose Chiplet Cloud, a chiplet-based ASIC AI-supercomputer architecture that optimizes total cost of ownership (TCO) per generated token for serving large generative language models to reduce the overall cost to deploy and run these applica- tions in the real world." So they are comparing actual implementations with a theoretical implementation. Never mind that they got the A100 figures wrong, they are still in the 'wouldn't it be nice if we had 'x'' stage. This looks like a paper whose sole purpose is to raise funds for a research project that will probably ultimately go nowhere and they needed a reason that looks good on paper to increase their chances of getting funded. A100 can already be had for $0.87/hour so even their theoretical advantage is under significant pressure and assuming they got everything else right by the time the project has run the market will have overtaken them. This is what usually happens to CPUs that are application specific.
- emmender 3y agotake 3 ideas that are hot: chiplets, cloud, and LLM - remix them into the title of a paper that describes a hypothetical machine.. academia playing catch up and trying to stay relevant in my cynical eye.
- fzliu 3y agoJust skimmed the paper. Seems to me like this paper wants to optimize transformer inference e2e, i.e. from ASIC level all the way to cloud. I'm not exactly convinced though, since all the results seem to be purely theoretical or simulated. I would've liked to see a prototype built across several FPGAs with clock speeds extrapolated for ASICs.
- londons_explore 3y agoIt seems fine to say "others have proved that this math makes a good LLM, we have designed an ASIC that can do this math fast, therefore we can make a good fast LLM"
- JonChesterfield 3y agoYes, but saying that shouldn't be mistaken for "we can make an asic that runs some model fast". There's a wide implementation void between the two.
- kraken12 3y agoYep, it's a research paper in comp arch, the initial proof-of-concept study before you go and spend real money on it.
- saturn99 3y agoI think FPGAs would be an awesome prototype but maybe too constricting in terms of resources? The extrapolation might be so far out to be just as accurate as their simulated model...
- SomeRndName11 3y agoThe problem with hardware solution, is lack of flexibility. LLM is not (yet) as established technology to warrant fixed in-silico solutions, compared to say GPUs.
- qwertox 3y agoSo one day we'll be buying LLM cartridges like we used to buy cartridges for the Atari.
- woah 3y agoYour ChatGPT6 cartridge is empty. Please replace your ChatGPT6 cartridge.
- DrNosferatu 3y agoAnd the performance on GPT4 and Falcon40B? If the design cannot serve models of this level, there will be no economic interest. And a comparison with Jim Keller's Tenstorrent AICloud?
- ilaksh 3y agoHow much have they optimized the software here? Is it tinygrad level optimization? Also does this lower total cost depend on SRAM being available for DRAM prices? What makes SRAM so much more expensive than DRAM?
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- skummetmaelk 3y agoSRAM uses multiple transistors and takes up a lot more space than DRAM so it is inherently more expensive because it needs more area. The advantage is that it is fast and doesn't need refreshing like DRAM does. You can also put it on the same die as your computation logic which is technically possible with DRAM, but kinda silly since you need an optimized process to get the best out of that. This process is then quite bad for high-speed logic.
- ilaksh 3y agoDo you think the price of including SRAM might be reduced somewhat if the big foundries optimize for including lots of SRAM in these types of ASICs?
- skummetmaelk 3y agoUnlikely. SRAM is already used heavily in ASICs today and lots of R&D goes, and has gone, into optimizing it already.
- gautamcgoel 3y agoCan anyone share how much SRAM their proposed ASIC actually has? I skimmed the paper but that number didn't jump out to me.
- punkgenius 3y ago~200MB to 1GB per ASIC, from Table 2 on page 10.
- btbuildem 3y agoWhen the LLM wave first burst into public consciousness, I hoped that people would find a way to repurpose all the crypto-mining hardware for this -- alas, a different set of problems.
- woah 3y agoASIC stands for Application Specific Integrated Circuit. So by definition, they cannot be repurposed.
- brucethemoose2 3y ago... Isn't this basically the Cerebras WS2? Each "die" has 40GB of SRAM, and they have a fast interconnect.
- punkgenius 3y agoThis seems to be more achievable and cost-effective than Cerebras. Some comments mention Cerebras cost millions for each 'die'.
- brucethemoose2 3y agoYes, that because Cerebras's "chip" is actially an entire wafer many GPUs would normally be carved out of. The extra stuff TSMC must do to pull that off are probably expensive... But I can't imagine it being, say, 10x more expensive than a wafer full of reticle sized dies (like Nvidia does). And thats setting aside the massive IO advantage of Cerebras's mega die.
- boredumb 3y agoWas expecting something else I suppose. Is this kind of stuff for potential future investors?
- saturn99 3y agoYep, pretty unique system. I imagine anything that runs LLMs better than we currently can would pique the interest of VCs.