5 ms·
This seems like a pretty bad paper. Their headline claim that they are 300x faster than an A100 at serving GPT-3 uses obviously wrong numbers for how fast A100s
by why_only_15 3y ago
This seems like a pretty bad paper. Their headline claim that they are 300x faster than an A100 at serving GPT-3 uses obviously wrong numbers for how fast A100s can run GPT-3. They seem to have misread the DeepSpeed Inference paper and claim that the best throughput on GPT-3 sized models was 18 tok/s, but if you look at figure 8 on page 11 of the paper [1], it shows that they are able to achieve ~74 teraflops on serving LM-175B (GPT-3), which is about 211 tok/s. To calculate TCO, we can say 211 tok/s = 760,000 tok/hour and A100s are about $1/hr, so TCO per 1k tokens using that paper's method is about $0.0013, much lower than the $0.02 that they claim, reducing the claimed TCO advantage from 94x to 6.2x. Combine this with the fact that the paper they used is from a year ago and there are more efficient inference methods than there were then and the speedup probably goes even lower, maybe to 3x. This is without even looking at the chip design itself, whose costs are probably far underestimated.
[1]: https://arxiv.org/pdf/2207.00032.pdf https://arxiv.org/pdf/2207.00032.pdf
- jacquesm 3y agoThe whole thing is imaginary: "In this paper, we propose Chiplet Cloud, a chiplet-based ASIC AI-supercomputer architecture that optimizes total cost of ownership (TCO) per generated token for serving large generative language models to reduce the overall cost to deploy and run these applica- tions in the real world." So they are comparing actual implementations with a theoretical implementation. Never mind that they got the A100 figures wrong, they are still in the 'wouldn't it be nice if we had 'x'' stage. This looks like a paper whose sole purpose is to raise funds for a research project that will probably ultimately go nowhere and they needed a reason that looks good on paper to increase their chances of getting funded. A100 can already be had for $0.87/hour so even their theoretical advantage is under significant pressure and assuming they got everything else right by the time the project has run the market will have overtaken them. This is what usually happens to CPUs that are application specific.
- emmender 3y agotake 3 ideas that are hot: chiplets, cloud, and LLM - remix them into the title of a paper that describes a hypothetical machine.. academia playing catch up and trying to stay relevant in my cynical eye.
- tmccrary55 3y agoUsing ChatGPT
- avereveard 3y agoI asked gpt for giggles and the comparison is much more thorough, it has written also power per watt improvements, benefits of denser packing, and sustainability of moving toward more energy efficient solution.
- nkko 3y agoI did the same exact mind exercise using ChatGPT but I haven't produced a paper out of the chat session.
- bhouston 3y agoThe $0.87/hour price you gave is theoretical and also we know any price in a paper for compute is wrong by the time of publication. Pragmatically the prices are closer to $2/hr according to this recent post here on Hacker News: https://llm-utils.org/Nvidia+H100+and+A100+GPUs+-+comparing+available+capacity+at+GPU+cloud+providers https://llm-utils.org/Nvidia+H100+and+A100+GPUs+-+comparing+... Although again prices change on a daily basis on spot providers.
- jacquesm 3y ago> The $0.87/hour price you gave is theoretical https://cloud.google.com/blog/products/compute/a2-vms-with-nvidia-a100-gpus-are-ga https://cloud.google.com/blog/products/compute/a2-vms-with-n... That's as close as I got to verifying that price.
- bhouston 3y agoAnother list that shows pricing both constant and spot. The best GCP spot price is $1.1, but Jarvis seems to say its spot for the 40GB A100 is $0.69: https://fullstackdeeplearning.com/cloud-gpus/ https://fullstackdeeplearning.com/cloud-gpus/ I feel there are more fair criticisms of that paper than its inclusion of the snapshot price of variable priced compute resource.
- jacquesm 3y agoSure, to me it more of an extra item than the main one but it is one that you can readily verify because most of the other claims are far more vague. If they're willing to fudge on that one then I have much less confidence in the rest of their claims.
- kraken12 3y agoSeems like HN comments have determined that the cost number is not fudged..
- kraken12 3y agoYeah, it is an architectural simulation study, this is what is usually done right at the beginning before resources are allocated to go deep on idea. So in that sense it is imaginary; but this is how new ideas get incubated.
- deleted 3y ago[deleted]
- beecafe 3y ago[dead]
- chessgecko 3y agoDoesn’t that graph have a toks/sec of 18? Or am I reading it wrong
- why_only_15 3y agoIt shows 18 tokens per second but that's how fast tokens are generated I think. The number of tokens generated is that times the batch size, which appears to be 12? The graph is quite unclear and I didn't feel like reading the paper more in-depth.
- kraken12 3y agoSeems to me the 18 tokens per second from [1] is the throughput and includes the batch size, so I don't think they misread the Deepspeed inference paper. So the chiplet ASIC supercomputer paper would seem to show a decent performance/TCO benefit. Of course, it's a first architectural study to illustrate the promise of the idea, lots more details to work out in a physical implementation and the final realized benefit is likely to be lower. But even a 3X is huge in this space.
- why_only_15 3y agoIf you look at the "metrics" section on Page 9, it says: > 2) Metrics: We use three performance metrics: (i) latency, i.e., end-to-end output generation time for a batch of input prompts, (ii) token throughput, i.e., tokens-per-second processed, and (iii) compute throughput, i.e., TFLOPS per GPU. This is somewhat confusing to me because at least two of these three definitions should be essentially the same thing, but I don't think there's any way to interpret their claim of ~74 teraflops achieved other than ~211 tokens/second of throughput. Put another way, 18 tokens per second is 2% flops utilization, which we are obviously capable of doing better than for bulk inference. 3x is not huge in this space because just using a 4090 instead of an A100 is a 5x gain.
- punkgenius 3y agoMaybe 74 tflops is the best they've achieved, but not all 16 GPUs can consistently hit that number? Just guessing.. The 211 tokens/sec throughput on GPU is just insane, it's even better than what TPU can do on PaLM 540B.
- punkgenius 3y agoLooks like figure 8 of paper [1] says it is 18 tokens/s
- saturn99 3y agoI think the costs are from the Moonwalk model, which is a pretty good reference for estimating costs, although it might be low if you use all Google engineers to build the HW. =P