10 ms·
Building a deep learning rig
- infogulch 3y agoI'm eyeing Tinybox as a deep learning rig. https://tinygrad.org/ https://tinygrad.org/ https://twitter.com/__tinygrad__/status/1760988080754856210 https://twitter.com/__tinygrad__/status/1760988080754856210
- Smith42 3y ago$15k!
- KeplerBoy 3y agoWhich is not unreasonable for that amount of hardware. You have to ask yourself if you want to drop that kind of money on consumer GPUs, which launched late 2022. But then again, with that kind of money you are stuck with consumer GPUs either way, unless you want to buy Ada workstation cards for 6k each and those are just 4090s with p2p memory enabled. Hardly worth the premium, if you don't absolutely have to have that.
- cyanydeez 3y agoI believe the ada workstation cards are typically 1slot cards which means you could build a 4gpu server from normal cases. most of the 4090 cards are 2-3 slot cards
- KeplerBoy 3y agoThe beefy workstation cards are 2 slots, but yeah the 4090 cards are usually 3.something slots, which is ridiculous. The few dual slot ones are water cooled.
- cyanydeez 3y agothe work station cards also run on 300 watts and looks like the 4090 goes to 450. so you are getting a better practical card for the price if you are making a mining type rig, then yeah, the extra price is wasting money. but if you wanted to build a normal machine, the workstation cards are the most reasonable choice for anything more than 2 gpus
- kkielhofner 3y agoI find it challenging to get my 4090s to consume more than 300 watts. There are also a lot of articles, benchmarks, etc around showing you can dramatically limit power while reducing perf by insignificant mounts (single digit %).
- justsomehnguy 3y ago> which means you could build a 4gpu server from normal cases. Only if you are already live near an airport and you are accustomed to the sounds of the lift off and flying away.
- treprinum 3y agoSure, if you want to waste your time on getting stuff working on AMD instead of spending it on actual model training...
- kkielhofner 3y agoThis. People complain about the "Nvidia tax". I don't like monopolies and I fully support the efforts of AMD, Intel, Apple, anyone to chip away at this. That said as-is with ROCm you will: - Absolutely burn hours/days/weeks getting many (most?) things to work at all. If you get it working you need to essentially "freeze" the configuration because an upgrade means do it all over again. - In the event you get it to work at all you'll realize performance is nowhere near the hardware specs. - Throw up your hands and go back to CUDA. Between what it takes to get ROCm to work and the performance issues the Nvidia tax becomes a dividend nearly instantly once you factor in human time, less-than-optimal performance, and opportunity cost. Nvidia says roughly 30% of their costs are on software. That's what you need to do to deliver something that's actually usable in the real world. With the "Nvidia tax" they're also reaping the benefit of the ~15 years they've been sinking resources into CUDA.
- dkjaudyeqooe 3y agoWow! It's incredible how Nvidia has created the dark voodoo magic, and how only they can deliver the strong juju for AI. How are the so incredibly smart and powerful!? I wonder if it has anything to do with the strategy they used in 3D graphics, where game developers ended up writing for NV drivers in order to maximise performance, and Nvidia abused their market position and used every trick in the book to make AMD cards run poorly. People complained about AMD driver quality, but the actual problem was that they were not NV drivers and AMD couldn't defeat their software moat. So here we are again, this time with AI. You'd think we'd learnt our lesson but instead people are fooled yet again and instead of understanding the importance of diversity and competition in the lifeblood of their art, myopia and amnesia is the order of the day. Tinygrad are doing god's work and I won't be giving Nvidia a single fucking cent of my money until the software is hardware neutral and there is real competition.
- tutfbhuf 3y agoThis is the new startup from George Hotz. I would like him to succeed, but I'm not so optimistic about their chances of selling a $15k box that is most likely less than $10k in parts. Most people would do much better by buying a second-hand 3090 or similar and connecting them into a rig.
- segmondy 3y agoNot necessarily, I'm not sure about AMD GPUs, but he tweeted that AMD supports linking all 6 together. If that's the case, then 6 of those XTX should crush 6 3090's. For us techies we definitely will decide to build vs buy. However businesses would definitely decide to buy vs build.
- tutfbhuf 3y agoI think businesses are more likley to rent and techies are more likely to build. So he is betting on a niche, techies who wants to buy a 15k device or companies who do not want to rent. Also companies who are willing to go the AMD GPU route instead of NVIDIA GPU, which as much better tooling and much more experts on the job market.
- downrightmike 3y agoHe keeps hopping from thing to thing
- varelse 3y ago[dead]
- lostmsu 3y agoAs I mentioned in comments to this post on twitter, you can beat this with a pretty regular $6000 2x4090 system on compute (but with less total VRAM).
- abra0 3y agoI was thinking of doing something similar, but I am a bit sceptical about how the economics on this works out. On vast.ai renting a 3x3090 rig is $0.6/hour. The electricity price of operating this in e.g. Germany is somewhere about $0.05/hour. If the OP paid 1700 EUR for the cards, the breakeven point would be around (haha) 3090 hours in, or ~128 days, assuming non-stop usage. It's probably cool to do that if you have a specific goal in mind, but to tinker around with LLMs and for unfocused exploration I'd advise folks to just rent.
- cyanydeez 3y agothe current economics is a low ball to get costumers. it's absolutely not going to be the market price once commercial interests have locked in their products. but if you're just goofing around and not planning to create anything production worthy, it's a great deal.
- whimsicalism 3y ago> the current economics is a low ball to get costumers. vast.ai is basically a clearinghouse. they are not doing some VC subsidy thing in general, community clouds are not suitable for commercial use.
- imiric 3y ago> On vast.ai renting a 3x3090 rig is $0.6/hour. The electricity price of operating this in e.g. Germany is somewhere about $0.05/hour. Are you factoring in the varying power usage in that electricity price? The electricity cost of operating locally will vary depending on the actual system usage. When idle, it should be much cheaper. Whereas in cloud hosts you pay the same price whether the system is in use or not. Plus with cloud hosts reliability is not guaranteed. Especially with vast.ai, where you're renting other people's home infrastructure. You might get good bandwidth and availability on one host, but when that host disappears, you should hope that you did a backup, which vast.ai charges for separately, and if so, you need to spend time restoring the backup to another, hopefully equally reliable host, which can take hours depending on the amount of data and bandwidth. I recently built an AI rig and went with 2x3090s, and am very happy with the setup. I evaluated vast.ai beforehand, and my local experience is much better, while my electricity bill is not much higher (also in EU).
- cyanydeez 3y agojust ordered a 15k thread ripper platform because it's the only way to cheaply maximize the pcie16x bottleneck. the mining rigs are neat because the space you need for consumer GPU is a big issue. those rigs need pcie riser slots that are also limited. looks like the primary value is the rig and the cards. they'll need another 1-2k for a thread ripper and then the riser slots.
- dijit 3y agoavailability is tight i think but check out the ampere altra stuff, they have an absurd number of pci’s lanes compared to AMD and especially intel, if you can suffer the ARM architecture. They also have some ML inference stuff on chip themselves.
- choppaface 3y agoBut then you need to deal with arm compile issues. A lot of common packages are available for arm, but x86 is still least likely to distract your development.
- segmondy 3y agoUnless you are training maximizing the PCIe lanes is truly overrated. You certainly don't want to be running at 1x speed. But 8x speed is enough with minimal impact. 8*3 = 32 lanes. Most CPUs can provide that. I'm running off a 2012 hp z820, that yields 3x16/1x8. So for anyone going for a build, don't throw money on CPUs. IMHO, GPU first, then your motherboard second (read the specs sheets), then CPU supported pcie lanes & Storage speed.
- kaycebasques 3y agoI really enjoy and am inspired by the idea that people like Dettmer (and probably this Samsja person) are the spiritual successors to homebrew hackers in the 70s and 80s. They have pretty intimate knowledge of many parts of the whole goddamn stack, from what's going on in each hardware component, to how to assemble all the components into a rig, up to all the software stuff: algorithms, data, orchestration, etc. Am also inspired by embedded developers for the same reason
- nirav72 3y agoThis is nice. I would’ve used one of those ETH mining cases that support multiple GPUs. Ebay has them $100-150 these days.
- whoisthemachine 3y agoI've been slowly expanding my HTPC/media server into a gaming server and box for running LLMs (and possibly diffusion models?) locally for playing around with. I think it's becoming clear that the future of LLM's will be local! My box has a Gigabyte B450M, Ryzen 2700X, 32GB RAM, Radeon 6700XT (for gaming/streaming to steam link on Linux), and an "old" Geforce GTX 1650 with a paltry 6GB of RAM for running models on. Currently it works nicely with smaller models on ollama :) and it's been fun to get it set up. Obviously, now that the software is running I could easily swap in a more modern NVidia card with little hassle! I've also been eyeing the b450 steel legend as a more capable board for expansion than the Gigabyte board, this article gives me some confidence that it is a solid board.
- Uehreka 3y ago> I just got my hands on a mining rig with 3 rtx 3090 founder edition for the modest sum of 1.7k euros. I would prefer a tutorial on how to do this.
- gigatexal 3y agoI thought this looked like a cryptocurrency miner. Seems the crypto to AI pivot is legit happening. And good. Would rather we boiled the oceans for something marginally more valuable than in-game tokens we traded for fiat funds in this video game we call life.
- neilv 3y agoFor large VRAM models, what about selling one of the 3090s, and putting the money towards an NVLink and a motherboard with two x16 PCIe slots (and preferably spaced so you don't need riser cables)?
- nick7376182 3y ago[dead]
- p1esk 3y agoWhy do you need x16 pcie slots if you can use nvlink?
- elorant 3y agoNVlink is to connect the cards to each other. To connect them to the board you need the PCie slots.
- p1esk 3y agoWe are talking about increasing the intercard bandwidth, assuming that’s a bottleneck. It can be done by either increasing pcie bandwidth, or using nvlink. If you use nvlink, increasing pcie does not provide any additional benefit because nvlink is much faster than pcie. p.s. the mobo (B450 Steel Legend) already has 2 pcie x16 slots, so the recommendation does not make sense to me.
- segmondy 3y agofull riser cables like they used doesn't impact performance. Hanging it off on open air frame IMO is better, keeps everything cooler, not just the GPU but the motherboard and surrounding components. With only 2 24gb GPU they are not going to be able to run larger models. You can't experiment with 70b models without offloading to CPU which is super slow. The best models are 70b+ models.
- ImprobableTruth 3y ago
- whimsicalism 3y agoI strongly, strongly suspect most people doing this are significantly short of the breakeven prices for transitioning from cloud 3090s. inb4 there are no cloud 3090s: yes there are, just not in formal datacenters
- soraki_soladead 3y agoIt's not always about cost. Sometimes the ergonomics of a local machine are nicer.
- smokeydoe 3y agoDoes anyone have any good recommendations for an epyc server grade motherboard that can use 3x3090? My current motherboard (strix trx40-xe) has memory issues now. 2 slots cause boot errors no matter what memory is inserted. I plan to sell the threadripper. Other option is to just swap out the current motherboard with a trx zenith extreme but I feel server grade would be better at this point after experiencing issues. Is supermicro worth it?
- KuriousCat 3y agoIt might not be the answer you are looking for, I would take a look at components published by System76/Lambda labs such as this to pick the one that would suit me: https://github.com/system76/thelio/blob/master/Thelio%20Common/Bill%20of%20Materials%20(BOM)/thelio-systems-bom.csv https://github.com/system76/thelio/blob/master/Thelio%20Comm...
- segmondy 3y agoIf you're just going to stick to 3 GPUs. Then a lot of consumer gaming motherboards would be more than sufficient. Checkout the z270, x99, x299. If you really want epyc go to ebay search for "gigabyte mz32-ar0 motherboard" Majority of them are going to come form China and they are all pretty much used. If you have plans to go even bigger then I say go for a new wrx80
- buildbot 3y agoI have this motherboard - a big downside is many of the PCIE slots will overhang into the RAM if used for a GPU. I can't use two channels in my current ML machine because of this, and I have single slot 4090s.
- devbug 3y agoH12SSL-i or H12SSL-NT ROMED8U-2T
- moondev 3y agoHeh I went the opposite direction from zenith to ASRock tnt Are you on the latest bios? I have had good results flashing the "E" bios in the TNT folder based on the FTP site mentioned here https://www.reddit.com/r/ASRock/s/GztBuD9INh https://www.reddit.com/r/ASRock/s/GztBuD9INh
- Yenrabbit 3y agoNote that they shared part two recently: https://samsja.github.io/blogs/rig/part_2/ https://samsja.github.io/blogs/rig/part_2/ For those talking about breakeven points and cheap cloud compute, you need to factor in the mental difference it makes running a test locally (which feels free) vs setting up a server and knowing you're paying per hour it's running. Even if the cost is low, I do different kinds of experiments knowing I'm not 'wasting money' every minute the GPU sits idle. Once something is working, then sure scaling up on cheap cloud compute makes sense. But it's really, really nice having local compute to get to that state.
- buildbot 3y agoLots of people really underestimate the impact of that mental state and the activation energy it creates towards doing experiments - having some local compute is essential!
- krallistic 3y agoThis. In the second article, the author touches on this a bit. With a local setup, I often think, "Might as well run that weird xyz experiment over night" (instead of idling) On a cloud setup, the opposite is often the case: "Do I really need that experiment or can I shut down the sever to save money?". Makes a huge difference over longer periods. For companies or if you just want to try a bit, then the cloud is a good option, but for (Ph.D.) researchers, etc., the frictionless local system is quite powerful.
- ummonk 3y agoI have the same attitude towards gym memberships - it really helps to know I can just go in for 30 minutes when I feel like it without worrying whether I’d be getting my money’s worth.
- deleted 3y ago[deleted]
- abra0 3y agoThat's a great point! I'd agree that just the extra emotional motivation from having your own thing is worth a ton. I get some distance down that way by having a large RAM no GPU box, so that things are slow but at least possible for random small one offs.
- 0x20cowboy 3y agoIf you would like to put Kubernetes on top of this kind of setup this repo is helpful https://github.com/robrohan/skoupidia https://github.com/robrohan/skoupidia The main benefit is you can shut off nodes entirely when not using them, and then when you turn them back on they just rejoin the cluster. It also helps managing different types of devices and workloads (tpu vs gpu vs cpu)
- 2OEH8eoCRo0 3y agoI love the idea of a "poor man's cluster" of hardware that I can continually add to. Old ereaders, phones, tablets, family laptops, everything. I'm not sure what I'd use it for.
- bick_nyers 3y agoSomewhat tangential question, but I'm wondering if anyone knows of a solution (or Google search terms for this): I have a 3U supermicro server chassis that I put an AM4 motherboard into, but I'm looking at upgrading the Mobo so that I can run ~6 3090s in it. I don't have enough physical PCIE slots/brackets in the chassis (7 expansion slots), so I either need to try to do some complicated liquid cooling setup to make the cards single slot (I don't want to do this), or I need to get a bunch of riser cables and mount the GPU above the chassis. Is there like a JBOD equivalent enclosure for PCIE cards? I don't really think I can run the risers out the back of the case, so I'll likely need to take off/modify the top panel somehow. What I'm picturing in my head is basically a 3U to 6U case conversion, but I'm trying to minimize cost (let's say $200 for the chassis/mount component) as well as not have to cut metal.
- choppaface 3y agoComino sells a 6x 4090 box as a product: https://www.comino.com/ https://www.comino.com/ They have single-slot GPU waterblocks but would want something like $400 or more each for them individually.
- ftufek 3y agoYou'll need something like EPYC/Xeon CPUs and motherboards which not only have many more PCIe lanes, but also allow bifurcation. Once you have that, you can get bifurcated risers and have many GPUs. And these risers use normal cables not the typical gamer pcie risers which are pretty hard to arrange. You won't get this for just $200 though. For the chassis, you could try a 4U rosewill like this: https://www.youtube.com/watch?v=ypn0jRHTsrQ https://www.youtube.com/watch?v=ypn0jRHTsrQ, not sure if 6 3090s would fit though. You're probably better off getting a mining chassis, it's easier to setup and cool down, also cheaper, unless you plan on putting them in a server rack.
- jeffybefffy519 3y agoAre m1/m2/m3 max mac's any good for this?
- downrightmike 3y agoWay slower than 1 gpu, at many times the cost. If you don't mind waiting minutes instead of seconds, macs are reasonable
- fragmede 3y agoIt depends on what you're trying to do, but I've got an M1, and doing inference with llama2-uncensored using Ollama, I get results within seconds.
- whywhywhywhy 3y agoDepends what you're doing, M1 Max is around a minute for 1 SDXL image and the machine feels like it's choking while it does it while a 3090 will do it in 9 seconds and doesn't feel like it's breaking a sweat. Llama definite a bit of a different story though.
- jeffybefffy519 3y agoIm more thinking about the training side because it could be compelling to buy a beefily specced m3 max if it can replace what a dedicated gpu rig could do and also be a daily driver.
- akasakahakada 3y agoJust sharing. 2 x RTX4090 workstation guide You can put two aircooled 4090 in the same ATX case if you do enough research. https://github.com/eul94458/Memo/blob/main/dual_rtx4090workstation_for_machine_learning_202401.md https://github.com/eul94458/Memo/blob/main/dual_rtx4090works...