9 ms·
This is really good, and I was really excited by it but then I read: > running on a single 8XA100 40GB node in 38 hours of training This is a $40-80k machine.
by arturventura 4y ago
This is really good, and I was really excited by it but then I read:
> running on a single 8XA100 40GB node in 38 hours of training
This is a $40-80k machine. Not a diss, but I would love to see an advance that would allow anyone with a high end computer to be able to improve on this model. Before that happens this whole field is going to be owned by big corporations.
- windexh8er 4y agoBut how often do you need to run this? You can run 8xA1000 on LambdaLabs [0] (no affiliation) for $8.80/hr. So you should be able to run the entire data set for less than $350. [0] https://lambdalabs.com/service/gpu-cloud#pricing https://lambdalabs.com/service/gpu-cloud#pricing
- throwawaymaths 4y agoThey are acknowledged at the bottom for supporting andrej's research!!
- ProjectArcturis 4y agoThat's to train it from scratch, though, right? If you preload the GPT2 weights you don't need to do this. You can just give it additional training on your texts.
- anilshanbhag 4y agoIf GPT-2 / nanoGPT needs this setup, just imagine what GPT3 / chatGPT needs!
- Gigachad 4y agoSupposedly even running the trained model for ChatGPT is extremely expensive unlike the image generators which can largely be run on a consumer device.
- anigbrowl 4y agoWell, he does include instructions for running it on a personal computer, which looks like what I'm gonna be doing next week. Besides the rental options discussed below these nvidia boxen don't look too big so either used ones will be available for cheap relatively soon, or you could just locate and liberate one in Promethean fashion.
- jph00 4y agoA couple of weeks ago a new paper came out that shows how to train a high quality language model on a single GPU in one day. https://arxiv.org/abs/2212.14034 https://arxiv.org/abs/2212.14034
- haldujai 4y agoIf you can’t fit the model on your resources you can leverage DeepSpeed’s ZeRO-offload which will let you train GPT2 on a single V100 (32gb). Alternatively, if you’re researching (with the caveat that you have to either publish, open source or share your results in a blog post) you can also get access to Google’s TPU research cloud which gives you a few v3-8s for 30 days (can’t do distributed training on devices but can run workloads in parallel). You can also ask nicely for a pod, I’ve been granted access to a v3-32 for 14 days pretty trivially which (if optimized) has more throughput than 8xA100 on transformer models. TPUs and moreso pods are a bit harder to work with and TF performs far better than PyTorch on them. https://www.deepspeed.ai/tutorials/zero-offload/ https://www.deepspeed.ai/tutorials/zero-offload/ https://medium.com/analytics-vidhya/googles-tpu-research-cloud-free-tpu-hardware-for-deep-learning-projects-7dfecd82b024 https://medium.com/analytics-vidhya/googles-tpu-research-clo...
- Tenoke 4y agoIt seems as likely as people being able to build big automaker level of cars just with tools in their garage. More compute is going to keep producing better results at least for LLMs.
- base698 4y agoYou can rent on AWS and other cloud providers.
- liquidk 4y agoThat is a key difference. You can’t easily and cheaply rent an auto factory, but you’re starting to be able to rent an LLM training factory once for a model where you can then more cheaply run inference on.
- krisoft 4y agoSo if I see it right that would be a p4d.24xlarge instance. Which goes for about $32.77 an hour nowadays so the total training would be about $1245. Not cheap, but certainly not a nation state budget. Edit: i just noticed lambda lab. It seems they ask $8.8 per hour for an instance of this caliber. That puts the total training cost around $334. I wonder how come it is that much cheaper.
- pavlov 4y agoI don't know if that's a blocker. Ordinary people commonly rent a $40k machine for 38 hours from companies like Avis and Hertz. If training a large model now costs the same as driving to visit grandma, that seems like a pretty good deal.
- Apofis 4y agoLet's not forget that rendering 3D Animations in 3DSMAX or Maya used to take days for a single frame for a complex scene, and months for a few minutes.
- deleted 4y ago[deleted]
- ofcourseyoudo 4y agoSimilarly maybe we should only let people rent a NanoGPT box if they are over 25 and they have to get collision insurance.
- jetrink 4y agoThat's a great comparison. For a real number, I just checked Runpod and you can rent a system with 8xA100 for $17/hr or ~$700 for 38 hours. Not cheap, but also pretty close to the cost of renting a premium vehicle for a few days. I've trained a few small models by renting an 1xA5000 system and that only costs $0.44/hr, which is perfect for learning and experimentation.
- willseth 4y agoThe good news is that, unlike vehicles, the rate for rented compute will continue to drop
- amelius 4y agoIt would be great if a tradeoff could be made, though. For example, train at 1/10th the speed for 1/10th of the cost. This could correspond to taking public transport in your analogy, and would bring this within reach of most students.
- aidos 4y agoI don’t know anything about this, but is that this instance type on AWS? p4d.24xlarge
- wongarsu 4y agoIt's a $33/hour machine on AWS, so about $1250 for one training run. Not cheap, but easily in the reach of startups and educational or research institutions. Edit: or about $340 if you get the 8xA100 instance from lambdalabs, in the realm of normal hobby spending
- bobbyi 4y agoIf you're doing something new/ custom (which you presumably are if you aren't using someone else's prebuilt model), it could take a lot of runs to figure out the best training data and finetune settings. (I assume. I've never worked with GPT, but have done similar work in other domains).
- weird-eye-issue 4y agoAfter training don't you have to keep it running if you want to use it?
- wongarsu 4y agoJust download the model and run it on something much smaller and cheaper. Bigger models like GPT-J are a bit of a pain to run, but GPT2-sized models run just fine on consumer GPUs.
- bilsbie 4y agoWhat’s required to run the model?
- wongarsu 4y agoThe biggest GPT2 (1.5B params) takes about 10GB VRAM, meaning it runs on a RTX 2080 TI, or the 12GB version of the RTX 3080
- renewiltord 4y agoWhat's the largest language model I can run on a 3090 with 24 GiB RAM?
- JustSomeNobody 4y agohttps://github.com/karpathy/nanoGPT#i-only-have-a-macbook https://github.com/karpathy/nanoGPT#i-only-have-a-macbook > This creates a much smaller Transformer (4 layers, 4 heads, 64 embedding size), runs only on CPU, does not torch.compile the model (torch seems to give an error if you try), only evaluates for one iteration so you can see the training loop at work immediately, and also makes sure the context length is much smaller (e.g. 64 tokens), and the batch size is reduced to 8. On my MacBook Air (M1) this takes about 400ms per iteration. The network is still pretty expensive because the current vocabulary is hard-coded to be the GPT-2 BPE encodings of vocab_size=50257. So the embeddings table and the last layer are still massive. In the future I may modify the code to support simple character-level encoding, in which case this would fly. (The required changes would actually be pretty minimal, TODO)
- dceddia 4y agoI was curious about how much this would be to rent, because definitely the cost of those servers is outside the budget! Lambda has 8xA100 40gb for $8.80/hr: https://lambdalabs.com/service/gpu-cloud#pricing https://lambdalabs.com/service/gpu-cloud#pricing
- kzrdude 4y agoHow are universities and colleges dealing with this kind of demand for computing power? It must be hard to be able to do some courses now.
- TrackerFF 4y agoAs far as research groups go - they get funds (project grants, donations, etc.) to purchase machines and parts, and then users have to timeshare them. These machines are pretty much crunching numbers 24/7, and your project will get appended to a queue.
- londons_explore 4y ago'group project'
- CuriouslyC 4y agoMost decently large colleges have been investing in HPC for a while, and started investing in GPU HPC around 2014. You'd be surprised what sort of school projects the compute budget exists for.
- r3trohack3r 4y agoI went to a smallish state university, even there we had our own HPC center and lab. We had a proper HPC (IIRC) 6 row data center across campus and we had a continuous budget available to me as an undergraduate research assistant for building beowulf clusters for the graduate programs to run assignments on. I once got an allowance to buy 15 raspberry pis to build an arm cluster.
- Tepix 4y agoIf you can fit the training into 24GB, a used RTX 3090 for $700-$800 seems like a good deal at the moment. They are about 45-65% as fast as the A100 according to https://bizon-tech.com/gpu-benchmarks/NVIDIA-RTX-3090-vs-NVIDIA-A100-40-GB-(PCIe)/579vs592 https://bizon-tech.com/gpu-benchmarks/NVIDIA-RTX-3090-vs-NVI... So if you buy two of these cards it will take 12-13 days instead of 38 hours but only require a $2500 PC. James Betker, who created tortoise TTS, built his own $15k machine with 8x RTX 3090 and trained the models with it. He now works for OpenAI…
- klaudioz 4y agoAny link to the 15k machine ?. Maybe it is cheaper now.
- Tepix 4y agoI think it was a DIY machine, those RTX 3090 have gotten cheaper for sure. From my experience, going beyond 4 GPUs is a pricey affair. See [§]. All but one model of the RTX3090 require at least 3 slots. If 4 GPUs connected via PCIe 4.0x16 are enough you can choose among various sRTX4 boards for 3000 series AMD Threadripper CPUs. [§] https://www.reddit.com/r/deeplearning/comments/tw0olq/comment/i3cuyks/?utm_source=share&utm_medium=web2x&context=3 https://www.reddit.com/r/deeplearning/comments/tw0olq/commen... Another useful URL: https://www.pugetsystems.com/labs/articles/Quad-GeForce-RTX-3090-in-a-desktop---Does-it-work-1935/ https://www.pugetsystems.com/labs/articles/Quad-GeForce-RTX-...
- Tepix 4y agoRecommended reading: https://timdettmers.com/2023/01/16/which-gpu-for-deep-learning/ https://timdettmers.com/2023/01/16/which-gpu-for-deep-learni... TL;DR: You probably don't need that expensive Threadripper because 2x PCIe 4.0 x16 will not be very beneficial. Go cheap, go 2x PCIe 4.0 x8.