7 ms·
Hey, author of the blog post here. It's mentioned in the blog post, but one of the intentions of this repo is that it's more of a "tutorial" than it is a librar
by chillee 3y ago
Hey, author of the blog post here. It's mentioned in the blog post, but one of the intentions of this repo is that it's more of a "tutorial" than it is a library/framework. My hope is that people will copy-paste and modify it for their own needs :)
Code can also be found here: https://github.com/pytorch-labs/gpt-fast https://github.com/pytorch-labs/gpt-fast
And a twitter thread summary here: https://twitter.com/cHHillee/status/1730293330213531844 https://twitter.com/cHHillee/status/1730293330213531844
- buildbot 3y agoGreat work and a really useful resource! Comprehensive guides on improving PyTorch performance are pretty hard to come by, and I learned a couple new tricks from this!
- ilaksh 3y agoWhat GPU was used when testing this? Is this faster than HuggingFace's Text Generation inference container?
- chillee 3y agoWe used an A100-80GB GPU. We didn't compare explicitly to Huggingface TGI but I think you should be able to compare the tokens/s achieved. One note is that this release is optimized for latency, while I think HF TGI might be more optimized for throughput.
- smith7018 3y agoGreat work! Do you know if it's possible to port this over to pytorch's Apple Silicon/MPS support?
- chillee 3y agoUnfortunately it's a little bit tricky today. The main issue is that we rely heavily on torch.compile + Triton for performance in this repo, and there isn't an Apple Silicon backend either for torch.compile or Triton. For example, there's an AMD backend for Triton (and it's also integrated into torch.compile), which is why we can mostly do the same optimizations on Nvidia and AMD GPUs. Ideally, there'd be an Apple Silicon backend for Triton, and then this repo would mostly work out of the box :)
- Dowwie 3y agoWhat kind of workstation would you build/buy for local GPT development with a budget of $3000? Is remote dev a viable alternative to local workstations?
- woodson 3y agoI’d go with a remote dev solution. Training/finetuning of large models requires much more resources anyway, so the GPUs in the local machine would be unused most of the time.
- leobg 3y agoNot OP, but I asked myself that same question two years ago. Then I looked at the energy prices in Germany and knew I had no chance against cloud GPUs. Maybe you live in a country with lower energy prices, like Bermuda (or any other country on earth), in which case this may not be as important to you. A side benefit of going cloud that you can pick and choose the right GPU for whatever project you’re working on, and you’re really just paying while you’re running them. Also, no hardware or Cuda drivers that may divert your attention.
- ftufek 3y agoLocal workstation is much cheaper in the long run. Even ignoring that, most of the development is running experiments. You're gonna be hesitant to run lots of experiments if they each cost money whereas when you pay upfront for the hardware, you're gonna have the incentive to fully utilize it with lots of experiments. I'd go with rtx 4090 and deal with memory limitation through software tricks. It's an underrated card that's as performant as cards that are magnitude pricier. It's great way to get started with that budget.
- Philpax 3y agoDepending on what you're doing, 2x used 3090s are the same price and offer you more VRAM. That's what I'm planning on doing, in any case - being able to run 70B LLMs entirely on the GPU is more useful than being able to run 34B faster.
- wolftickets 3y agoJust wanted to share, the charts and gifs are exceptionally well done. Informative, concise, and easy to read.
- chillee 3y agoThanks! I've also written a couple other things along a similar vein you might like at https://horace.io/writing.html https://horace.io/writing.html (particularly https://horace.io/brrr_intro.html https://horace.io/brrr_intro.html) and also some of the things I've tweeted: https://twitter.com/cHHillee/highlights https://twitter.com/cHHillee/highlights
- toxik 3y agoFantastic article, great job. One small note: eke and eek are not the same word.
- _giorgio_ 3y agoWhat's the difference between, say, Karpathy's nanoGPT and your GPT implementation? Is it just a difference in speed, or are there some new theory details to learn? Thanks for sharing your work.
- chillee 3y agoJust a difference in speed. This repo is primarily showing how you can get really good inference perf with just native pytorch.
- kartoolOz 3y agoHi, Thanks for Open sourcing the code! I was trying to reuse the code especially the dynamic quantization per channel (int8 on gpu) but couldn't get it to work, i also checked out torchao package but it looks like it has dependency on the nightly channel and SAM's dynamic implementation with triton has other issues, is there any clean implementation of int8 dynamic post-training quantization that you can point too ?
- chillee 3y agoWhat’s the issue with getting int8 dynamic quantization to work? As in, you’re unable to get it to quantize or to run with speedups?