15 ms·
Was my $48K GPU server worth it?
- timw4mail 5mo agoAnd here I felt like I was wasting money on an Intel B70 to run LLMs locally.
- doctorpangloss 5mo ago> Because of this I got a motherboard with slow GPU interconnect. It’s good for running many small experiments in parallel (which is my main use case) but horrible for any models split across gpus. :( you paid a professional pc builder and you weren't told this?
- ginko 5mo agoDon't those Ada 6000 GPUs support NVLink? I think I can even see the cover for the connectors in OP's pic. edit: Hm, finding mixed information online on whether that's still supported or not. Apparently it was removed in workstation GPUs.
- mciancia 5mo agoNope, they don't support it. And afair even if they did, you would be limited to connecting only in pairs, not all 6 together
- ryandrake 5mo agoHonestly, I made the same mistake when I added a GPU to my (not $48K) existing homelab. I got a Ada 4000 for its slim form factor and low wattage, but realize after I bought it that it does not support NVLink, so I can't really effectively double it up later if I wanted to. Live and learn. I suppose you might research that a little before blowing that much money though LOL :)
- deleted 5mo ago[deleted]
- CamperBob2 5mo agoConsumer motherboards can still make sense even if you leave some performance on the table. Running an actual 8x GPU server is not something you'd want to do in an apartment. Imagine the old Lucasfilm "THX" trailer where an unearthly-sounding foghorn whine rises to a sweeping crescendo at reference level, only without the decay at the end. At the time he put this rig together, there weren't a lot of open-weight LLMs that could run well on 6x48=288 GB, so it probably wasn't a huge loss. There still aren't, really. Right now I'm in the process of cramming Blackwell cards into an old DDR4-based Milan server, where the important thing is to be able to run large models at all. The GPU fans alone burn over 400 watts at full throttle.
- storus 5mo agoDid you think about Max-Q cards? 300W and they aren't that noisy either, 14% lower perf than non-Max-Q card.
- CamperBob2 5mo agoThat was an option, but having decided on a true server chassis for other reasons, it made sense to use server-edition cards to take advantage of all those fans. I downclock them to 300W anyway for longevity, but it's nice to have the option to go to 600W if needed. The server is going to live in the garage, so I'm not that concerned with noise. But I had no idea what to expect when I flipped the switch for the first time. It sounds like something out of the Book of Revelation. No way, no how could something like this be used in an inhabited area.
- mciancia 5mo agoI wonder why using 2 PSUs resulted in having slower interconnect. There is no specs in this blogpost regarding cpu/motherboard choice, but if you go with threadripper pro they have 128 pci-e lanes for some time now, so using all GPUs at full speed shouldn't be a problem
- thecatmak 5mo ago[dead]
- m-hodges 5mo agowhat is a "professional pc builder" in 2026
- ok_dad 5mo agoA guy on Facebook with more confidence and better insurance
- zozbot234 5mo agoIf you split models using pipeline/layer parallelism you don't have to care about a slow interconnect, you're just slowed down a lot when running a single inference at a time as opposed to a fully pipelined minibatch. But tensor parallelism requires much faster interconnects than you could get in your average server, so I'm not sure that a different motherboard would help all that much.
- shout5 5mo ago> paid a professional pc builder They did not. That's a mining rig not a workstation. It's visible from the photo and the chart showing multiple failures over a short period of time including the risers -- which are visibly very low quality -- failing twice. You have 50K, you call a real expert like Puget Systems or Digital Storm.
- gosub100 5mo agoIt doesn't cover risk. If one or more gpus dies, who pays for it? If you rent, you are guaranteed to be insulated from this risk. But owning, you might not have the best return policy from the vendor. And if you are actually at fault for breaking it, they have every right to deny a return. Or if your apartment is burglarized or catches fire (possibly from overloading the circuit) you are out the entire investment.
- 0xbadcafebee 5mo agoAlso a lightning strike or surge from the electric utility could fry the whole rig. Proper protection costs thousands, and even then it's not guaranteed to protect everything
- mschuster91 5mo ago> Proper protection costs thousands Frankly that's something a landlord should provide. And there's insurance against losses from electrical issues.
- beachy 5mo agoWhy should a domestic landlord provide you with data center-level power protection instead of just the normal household utility connection?
- mschuster91 5mo agoI'm talking about standard surge protectors. Properly installed they are enough except for direct lightning strikes, these will fry everything. But unfortunately, even in code-obsessed Germany landlords are not required to retrofit SPDs.
- 0xbadcafebee 5mo agoTo protect a large electrical device investment, you would want an EMP shield whole-home SPD, in addition to an SPD right at the electrical device. The first one shields exterior surges (including non-terrestrial), but the second shields against internal surges. And yeah lightning will blast through both of them. So the best bet is probably a lightning strike detector combined with renters insurance.
- tombert 5mo agoI have four old 24gb Nvidia cards. They're not great but they're not useless either. The problem is that I haven't really figured out a good way to actually use them. Genuine question; would anyone here recommend any specific motherboard to best utilize these cards?
- mciancia 5mo agoDepends what you want to do and which cards you have, but usually going with any older (3rd gen+) threadripper pro setup will give you a lot of pcie lanes. I myself run with gigabyte trx40 aorus xtreme, but since it's regular threadripper (not pro) with 4 GPUs 2 of them will run at x16 and two of them at x8 speeds
- throwawaytea 5mo agoYou could ask AI and get pretty far reading the answer.
- tombert 5mo agoI know. But this is a forum filled with technical professionals and I would like to get actual opinions from actual humans. AI is cool but it's not going to have all the good and bad experiences that humans have had with different motherboards.
- throwawaytea 5mo agoActually that's the best part of AI. It has access to experience with way more than the select sample size here.
- tombert 5mo agoI'm not entirely sure what your point is here; me asking for humans to give an opinion does not preclude me from also asking AI.
- 5mo ago
- hasteg 5mo agoJust curious OP (if you're the one posting) -- what do you mean by independent researcher? What are you researching and are you making $$ from it or are you living off previous built up savings? Seems like an interesting path. What research have you looked into so far?
- exceptione 5mo agoI am not the author, but he has been training/tuning? a model that produces text that mimics the source material in a more natural way. So getting the LLMs to produce less bland and boring LLMisms, according to the following up blog post.
- hsuduebc2 5mo agociting from the article: "I spent a long time trying high risk/high reward experiments and failing. But now I have something good. I’ve solved a major problem with LLMs. And I’m launching next Monday so we will soon see if it’s actually a breakthrough or just LLM psychosis " Maybe ai companies today have some bounty program?
- daemonologist 5mo agoThey have a subsequent post (from Monday) about what they've been working on: https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distribution-fine-tuning/ https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distri... (I would assume they haven't made a lot of $ off of this, if nothing else because they've only just put out that post and demo. They do seem to have produced a model that doesn't sound very LLM-y to my ear, though it also seems rather weak for its size.)
- bityard 5mo agoShallow take: They made an LLM that uses fewer emdashes. Cynical take: They made an LLM that can bypass existing AI slop detectors. Realistic take: They found a research problem they found interesting, dumped a bunch of capital and sweat equity into and (claimed to have, at least) found a solution. Neat!
- freediddy 5mo agoIn the last year, I have bought an M3 Ultra Mac Studio with 512 GB, a Macbook Pro M5 MAX with 128 GB and an RTX 6000 Pro. I have spent around $25k so far, not including electricity. I figured worst case scenario I can sell them in the next year and only take a haircut as opposed to losing my entire investment. In comparison to just spending for tokens, the tokens would have been much cheaper and much much faster. I've been running against Gemma4:31b, Qwen3.5 and 3.6, and getting local LLMs to solve AMC 8/10 math questions and it's about 10-100x slower than just doing it online. When I tried it with ChatGPT late last year, it took about one night and $25 to solve about 1000 questions. Using my RTX 6000 and M3 Ultra and Gemma4:31b on both, it answered about 40 questions in 7 hours and I haven't checked how good the answer is yet. At 800 watts (600 for RTX and 200 for M3 Ultra) and running for 7 hours, it solved around 40 questions. At the very least I'm going to try to sell my M3 Ultra if I can find a reliable place to sell it without getting ripped off by scammers.
- ahmadyan 5mo agoIf you are in the bay area, i'm happy to buy that M3 Ultra from you, i've been unsuccessfully looking for one and can't find any.
- jon-wood 5mo agoI’m not usually one to ask this because learning to do a thing can be fun, but why exactly have you spent 25 thousand dollars on getting an LLM someone else made to answer maths exam questions?
- freediddy 5mo agoIt's just a project I'm working on. I'm working on projects where AIs are processing and classifying large amounts of data that would be a lot of work for humans to do.
- wutwutwat 5mo agoI think of LLMs as being well equipped for handling dynamic data or adapting to unforeseen circumstances well (random code requests, website's ever changing layouts, typos, non-standard formatting in docs, groking out important info, etc), but math problems are be definition a very specific set of instructions to run, so is the overhead and "thinking" aspect of a LLM/AI even needed here? I'm genuinely curious, btw, I'm not asking sarcastically. Can't these math problems just be yanked from some test file and rapid fired directly at a gpu/compute unit?
- Aurornis 5mo agoThis is a difficult calculation to make because you wouldn't rent time on the exact same system in the cloud. Depending on what you're running, a bigger server with better inter-GPU interconnects in the cloud might complete the task so much faster that the additional per-hour expense is more than covered.
- peheje 5mo agoAgreed. And the gained time either goes toward 1) more experiments, or 2) leisure, which makes you sharper in the lab and happier overall. Not sure the "I saved $17,000 so far" framing is the most useful way to look at it, but it's a cool project and I love that people are doing this kind of thing.
- twobitshifter 5mo agoRight, you can rent from a v100 from llama cloud for $0.79/hr. An h100 is $3.99 /hr. $48000 is equal to 12000 hours of renting an h100, which is about as long as you’d spend at your job for 6 years!
- jmyeet 5mo agoSo some things have changed since this rig was first built (2024). The most relevant is that $6800 RTX 6000 Ada 48GB has arguably been supplanted by the $9500 RTX 6000 Pro 96GB. The Ada has a memory bandwidth of 960GB/s. The Pro has 1.8TB/s and about 40-50% better performance so is at least equivalent in processing power, much better in memory bandwidth (important for inference) and can hold larger models on a single card. I've considered buying a rig with 1-2 6000 Pros for similar reasons but I want to see what happens with this year's Mac Studios with a likely M5 Ultra. Macs have a shared memory architecture whereas NVidia segments the market based on max memory where the biggest consumer card (RTX 5090) has 32GB of VRAM but still excellent memory bandwidth (1.8TB/s). A RTX 5090 rig will still trounce a Mac Studio seems to be the conventional wisdom. Despite being able to hold larger models and being able to chain Mac Studios on TB5, their lower memory bandwidth (~900GB/s) and lower overall GFLOPS mean they still come out behind. That being said, the current Mac Studios are relatively long in the tooth, being released in 2024. I'm still not sure any of this is really wroth it because things are still changing so fast. I think there's a decent chance of a number of large AI companies going bust in the next 2-3 years such that you'll be able to buy enterprise AI hardware at cents on the dollar, a bit like how Google bought data centers in the post-dot-com crash. But anyway, nowadays I'd be looking at the RTX 6000 Pro as the sweet spot, having anywhere from 1-4 in a single server. The electricial issues the author mentions are interesting. I hadn't really thought about the max amperage on a residential circuit. In a DC, these would typically operate on three phase power and much higher overall amperage. I wonder if there's a device you can buy that can combine multiple residential circuits into a single power source for a server this power hungry?
- freediddy 5mo agoI have the Macbook M5 MAX with 128 GB of RAM. I put its performance at roughly equivalent to the RTX 5070 Ti. The M3 Ultra 512 GB for me is about half the performance of the RTX 5070 Ti but obviously it has the ability to do more because of the increased memory. I don't think anything compares to the nVidia chips at all.
- nextos 5mo agoI am also considering to buy 3-4x RTX 6000 Pro 96GB plus some Ryzen workstation with a grant. Is this the best general-purpose choice as of 2026 with $50k for training, fine-tuning and running large open models?
- 0xbadcafebee 5mo agoSo the answer is: "TBD if I can actually make money to pay this back"
- Quarrel 5mo agoIf nothing else, rosmine's DFT [1], which is what they were working on with this setup, seems like a worthwhile investigation. While I'm skeptical that there is much of a moat, at least for the large players, it should at least hopefully set rosmine up with for the next job :) It does seem to fix the current biggest issues with using LLMs for writing at various publishers. If you're The Economist, you have a very specific house style and you have a decent corpus of articles written in that style. At least on my reading of it, rosmine can use DFT to get a model to closely match its outputs, in terms of the language quirks that are generated, to that of the corpus it is fine tuned on. ie it will very much match the house style, particularly as it is used in writing, vs giving a system prompt to an LLM that has some Economist articles in its vast training set, and telling it to write in that style- it will do an ok job, but still exhibit LLM language quirks despite itself. Even if you feed it the specific "style guide" that they give their authors, I dare say the reality of their writing is the best place to learn, and it sounds like DFT can ground the writing of a model in a specific corpus like that. [1]: https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distribution-fine-tuning/ https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distri...
- vidarh 5mo agoGiving an LLM samples and tell it to apply the style in the sample works a lot better than just telling it to copy a style it may have seen, or a list of rules. They do it well enough that it'd take really good output to beat.
- Quarrel 5mo agoThey really don't. If your goal is to say, write science fiction, their reversion to classic LLM-isms, is really distracting and is what makes people say from a glance that it was written by an LLM. You basically can't use them at the moment in any real "natural" long-form writing. Everyone will call "slop" pretty quickly on the current frontier models. Rosmin's DFT paper is worth a read.
- jameson 5mo agoThe idea is similar to maintaining on-prem vs cloud Cloud is optimized for development velocity but its nature of high margin business eventually makes on-prem more promising It could be too late but it might be worth looking into tax saving if you have a business. Depreciation of asset is a loss and may deduct your income. (I'm NOT a tax expert)
- deleted 5mo ago[deleted]
- jmyeet 5mo agoCloud servers have cheaper electricity, the scale of industrial-level cooling, no issues for you (as a user) with hardware failure (ie you just use a different server; it's not your problem) and can amortize their cost by running 24x7. I've seen H100 computer hours for as little as $2. As the author notes, there are also electrical/wiring issues that cap how much compute gear you can run in a space not designed for it. I suspect a standard 20A 110V circuit can probably handle 2x RTX 6000 Pros. 15A probably can but that requires more research. Anything more than that and you're using multiple circuits, which has issues, or you need an upgraded circuit (eg 40A 240V) with all that entails (eg heavier duty cables, custom plug, etc).
- CamperBob2 5mo agoI suspect a standard 20A 110V circuit can probably handle 2x RTX 6000 Pros. 15A probably can but that requires more research. During initial setup of the server I am putting together, I found that a machine with 4x Blackwell cards derated to 300W can get by on a single 120V 20A circuit. It's tight but doable. A lot depends on the power supply. I don't think it's a great idea to run 4 high-power GPUs on a single ATX-style PSU, even a beefy 1600W job. The other questionable part is whether all four cards can temporarily spike at full power during boot, before the wattage limit is applied by the OS. Some accounts say this is possible, and if so it could shut down the party in a hurry. But I didn't see any misbehavior when I tried it.
- christina97 5mo ago
- datadrivenangel 5mo agoI did the math at least on a Macbook pro, and for inference it's definitely not worth it. - https://www.williamangel.net/blog/2026/05/17/offline-llm-energy-use.html https://www.williamangel.net/blog/2026/05/17/offline-llm-ene... - Discussion: https://news.ycombinator.com/item?id=48168198 https://news.ycombinator.com/item?id=48168198
- jmyeet 5mo agoIt's comparing laptops to dedicated GPUs in a server environment. The best comparison would be the Mac Studio but the current release is almost 2 years old at this point. We'll see what a likely M5 Ultra Mac Studio looks like, probably in Q3 this year. But yes, for pure inference, the M5 Max Macbook Pros probably aren't there yet. They have other utility though of course. And you can get 64GB and 128GB MBPs at a discount. Micro Center currently will let you buy a 64GB M5 Max MBP for under $4k currently, for example.
- joefourier 5mo agoWhy didn't you take into account batching, input tokens, different costs of electricity, and the fact that a laptop can still hold a decent % of its resale value, and is useful for many other tasks than running an LLM?
- bigyabai 5mo ago> Why didn't you take into account [...] the fact that a laptop can still hold a decent % of its resale value, and is useful for many other tasks than running an LLM? Because that wasn't what they claimed to research? >> for inference it's definitely not worth it. It's entirely fine if you enjoy local LLMs on your computer, there are people doing horribly inefficient inference on smartphones now. But for pure inference tasks, it's pretty obvious why M5s and Mac Studios aren't replacing TPUs and GPUs.
- joefourier 5mo agoWho is going to buy a $4299 M5 Max MBP with 64GB of RAM just to run Gemma 4 31b? Firstly you don't need 64GB for that model. Secondly if you want a machine that sits in the corner and does nothing but LLM inference, you don't buy a MacBook Pro, you buy some GPUs which are going to cost you a fraction of that (~$1k for ~64GB of VRAM is possible). The people buying Apple Silicon for inference general aim for the Mac Studios with enormous amounts of RAM (128-512GB), to run very large models. The idea is obviously to be running the LLM on your work laptop. As a developer I'd need a laptop with 24GB of RAM for work anyway, and 48GB, which is enough for a very good quant of Gemini, is just $400 extra.
- pelasaco 5mo agoout of curiosity, did you check how much would cost to rent a cage in a colocation space? Having to power your computer from two different outlets sounds wild..
- forsalebypwner 5mo agothe very last line of the article: "If I were to do this again, I wouldn’t do a custom build like this. I would buy a standard datacenter server and rent space in a colocation center. But then I would miss saying Hi to grumbl once in a while."
- pelasaco 5mo agoYes, i mean, he could rent a cage and run grumbl it there. It doesn't have to be a standard datacenter server, even though a standard datacenter server would be better and cheaper.
- kube-system 5mo agoA cage[0] is ~100x larger than what you need to host a single server. Many data centers will colocate by the rack unit. At others you can get a quarter or half cabinet[1]. Even at the very largest enterprise datacenters you can colocate a single cabinet. [0]: https://static.cisco-eagle.com/images/category/WireCrafters/Datacenter-Security-Cages.jpg https://static.cisco-eagle.com/images/category/WireCrafters/... [1]: https://www.edpeurope.com/wp-content/uploads/EDP-3-Compartment-Co-Lo-Rack-800x800.jpg https://www.edpeurope.com/wp-content/uploads/EDP-3-Compartme...
- pelasaco 5mo ago> A cage[0] is ~100x larger than what you need to host a single server. Yup, but i was assuming that he wanted to experiment building gpu rigs. For sure standard GPU servers are cheaper and easy to maintain. I have two lenovos, bought them used, already EOL.. was cheap and better than any custom gpu rig.. but i was pragmatic, because my goal was to put it in production, and not to research...
- dekhn 5mo agoI can't imagine spending $48K on a home GPU server, but I did just splurge and buy a PC with an RTX 5090, specifically to hold the largest models you can fit in 32GB. It's a top of the line PC with water cooled high end CPUs, 64GB RAM, RTX 5090 for $5K. To me the jury is still out whether this was a worthwhile investment, but I do expect to use this machine for a decade. I don't run it at 100% power (it's mostly idle, except for times when I'm training or doing batch inference). It has the nice property of being blackwell generation, similar to the machines we use at work. It just scares me to own a box that is $48K in my house, especially if it breaks, or gets stolen.
- throwatdem12311 5mo agoNot even a single mention of gaming. No wonder gamers hate AI bros.
- dekhn 5mo agoI have a second computer with an RTX 4090 for gaming (running Windows). I also used the new RTX 5090 running Linux to evaluate whether Proton/Wine allow me to run Windows games on linux (yes, it works, but the compatibility and frame rate issues make me stick to native Windows for now).
- fortyseven 5mo agoI wonder what's going wrong there? Personally I found compatibility and performance on Linux to be extremely good. And just keeps getting better. And that's not even just me, that's all kinds of benchmarks out there. Sorry to hear that. : ' (
- dekhn 5mo agoNo idea. I agree that in principle I should have close to the same performance on Linux. I just didn't want to spend a bunch of time customizing configs and updating software so I could reach parity with Windows when I had two computers.
- deleted 5mo ago[deleted]
- amarant 5mo agoThe research that's presented in another article on the same site is way more interesting than the betteridges law article linked here. It'll be very useful in my own latest project if this research is incorporated into some model I can rent by the token!
- janalsncm 5mo ago(For reference I’m talking about the DFT post from the same blog.) I love that ML is still in the “gentleman researcher” stage where relatively small amounts of startup capital can buy a ticket into frontier research. For a lot of research questions 6 GPUs is even overkill. It’s one of the reasons I’m skeptical of the “trillion dollar supercluster” idea [0]. I think what we need is more reasonably smart people investigating medium-sized problems. A “GPU middle class” you might say. [0] https://situational-awareness.ai/racing-to-the-trillion-dollar-cluster/ https://situational-awareness.ai/racing-to-the-trillion-doll...
- rosmine 5mo agoI agree :) Also, I heard Teknium trained the original Hermes model on 2x 4090. You can do a lot with a little compute
- gwbas1c 5mo agoFYI: If you're in a similar situation, think very carefully before you build your own. The $17000 might sound like a lot; but when you take into account your time and risk tolerance, renting might be a much better solution.
- forsalebypwner 5mo agoI think their retrospective at the end of the article is grounded and logical: "If I were to do this again, I wouldn’t do a custom build like this. I would buy a standard datacenter server and rent space in a colocation center" I'm sure there are use cases when renting makes sense, but it can get crazy expensive really fast if you're not careful.
- NicoHartmann 5mo agoJensen Huang said 'the more you buy, the more you save,' and you actually took it personally.
- irusensei 5mo agoI've hit command+f and then looked for this.
- NicoHartmann 5mo agoGlad I could fulfill your search query. Doing my part for the SEO of this comment section.
- rosmine 5mo agoHi! Thank you so much for posting this! I got back luck/timing when I tried, so happy it made it to the front page! (I am the author)
- 4chandaily 5mo agoI did this with used parts and cheaper consumer cards (3090s) and did much of the same calculations. I found it was way cheaper for me as well. The main advantage, however, is that the friction of "this is going to cost me in tokens to even try" goes away. I was so much more willing to take chances and try new things on my own hardware than I would have been if I were paying API costs. I feel like this point isn't made clearly enough by those of us who run these absurd self-hosted inference systems. Thanks for the write up, was a fun read. I spent an order of magnitude less, but I could relate to your story from beginning to end. Epyc (Milan), 512gb ram, 4x 3090
- bjt12345 5mo agoYou kind of bury the lede in that Article, it's a good article, well done getting interest in your work. Will you now be selling these GPUs for a profit?
- LandoCalrissian 5mo ago'If you google “plugging a PC into multiple outlets”, you get lots of warnings that if you even consider such a setup you will instantly burst into flames. So I hired a professional PC builder make sure it was safe.' Not really sure how that makes it safe but OK!
- 4chandaily 5mo agoI read this as; the "professional PC builder" would carry some sort of insurance. So it isn't really "safer", but if something goes wrong, the investment is (potentially) still safe. Just an assumption, though!
- kccqzy 5mo agoProbably means hiring someone who has more knowledge about PSUs and especially about having two simultaneous PSUs. There are questions like: when you press the power button how do the two PSUs turn up and in what sequence? How do you deal with the PWR_OK signal? What if there are voltage differences between the two PSUs? What about power backfeeding?
- mrandish 5mo agoI guess it was supposed to be a humorous aside, but it wasn't actually helpful because the relevant issue is when you pull more total amps from a single circuit than it's fused for (usually 15 or 20 amps in U.S. residences). The failure mode is usually tripping the circuit breaker. That issue can often be addressed fairly easily by splitting the power draw between two adjacent circuits. You can have an electrician do it permanently or temporarily DIY it with an appropriately rated extension cord. The real issue was OP was in an apartment at the time so an electrician would have been difficult. I assume they decided to just have a system integrator build it because they didn't want to figure out how to segment and route the power rails in a dual power supply system, but it's not exactly rocket science. Problems are often more due to choosing power supplies that aren't up to their claimed spec, not pre-testing them under load or using incorrect or under-spec cables.
- ttshaw1 5mo ago
- m463 5mo agoOther things people spend "too much money" on: - muscle cars, with all the stuff, driven occasionally. - boats, that don't get taken out much - gamer x, where x=system or laptop or keyboard or mouse or desk or glasses or mousepad or speakers or ... usually with "> too much RGB" - children $48k for something constructive even if ai related? no problem, refreshing even.
- dylan604 5mo agoOne of these things is not like the others. If you don't spend the money on one of them, you can get a visit from government officials that might decide to take that "item" from you. You'd also be a worthless human to spend that money on the other 3 while not on the one.
- m463 5mo agoI (probably not obviously) meant the "too much" part, where kids fail to grow/launch when (optional) things are given to them too easily. I didn't mean food, shelter, medical, education, lego, others-where-required-by-law.
- BLKNSLVR 5mo agoI have profited richly from my children, but not in a monetary sense. However much it has cost me monetarily, it has repaid itself ten times over in value to my very soul.
- rib3ye 5mo ago> I thought that I could not get a standard datacenter server because my apartment wouldn’t let me upgrade the circuits, so I needed to have 2 power supplies plugged into different circuits. Why didn't they just put a higher amp breaker in the box?
- chaidhat 5mo agoIt is unsafe for wires to be handling higher power than it was rated cause the wires act like very low ohm resistors. At some high enough I, you’re still gonna be generating power P=I^2R which is mainly thermal and melt the wires.
- pwg 5mo ago> Why didn't they just put a higher amp breaker in the box? 1) note the word "apartment" -- they rent, not own, and doing so not only would likely be illegal, but might also get them kicked out of the apartment. 2) Unless the wiring on the circuit drop, and all the end points are rated to handle the higher current, doing so would be an electrical code violation (and therefore trip into that "illegal" arena that might result in getting kicked out of the apartment). Most residences are wired using the minimum size wire rated for the installed breaker (because doing so saves costs). So a 15a breaker in the box would mean 14gauge (the US NEC minimum size for 15a circuits) wiring in the walls and 15a rated outlets/switches. Installing a 20a breaker in the box would be a code violation, and in many jurisdictions also illegal. And all the above is without considering that installing a 20a breaker on wires rated for 15a increases the fire risk tremendously if those wires are now asked to actually carry 20a for any length of time.
- hmokiguess 5mo agoStuff like this + OpenClaw with Mac Minis a while back is sort of exposing a probable local AI flywheel waiting to happen. Someone needs to solve proper distribution of packaged GPUs with some Tesla-like wall connector for a consumer grade box that is plug and play. Maybe John Ternus ends up doing that at Apple since they sit closer to this consumer profile.
- cpard 5mo agoUPDATE: Launch was a success! 400K+ views, and multiple companies reached to use my IP. Read more here It seems that he managed to get what he wanted from the hardware and I'm happy for them. He said something interesting at the beginning of his post, he compared the cost of the hardware to the cost of his time based on his FAANG salary. Which is an interesting way to think of this, but the rest of the article didn't make me understand if at the end he did save money/time based compared to just rend on the cloud. Also, outside of the power cost, hardware has other costs too, you need to operate it, maintain it, set it up, etc. all that require time. I mean, even the process of figuring out if it had a good enough ROI compared to cloud, takes from your time (collecting data, analyzing data, etc etc).
- utopiah 5mo agoDoubt it, feels basically like just an ad to get attention "Oh look, that's where the magic happens" vs running their code on existing infrastructure and thus just showing the results, like everybody else. This "feels" more "tangible".
- MagicMoonlight 5mo ago[flagged]
- caymanjim 5mo agoThis article appears to lack any reason for "needing" this beast, or any real comparison with alternatives, both of which are required to answer the question posed in the title. It's a summary of how much they spent and some light anecdotal comparison to what they might have spent on cloud services, but clearly they didn't do an exhaustive hunt for value. The real question is whether or not they could have done whatever it is they did with less hardware. Is there a business idea here that could have been proven on cheaper hardware that could be upgraded as demand increased? Is the expected ROI there based on future earnings? Absent any indication that this was needed in the first place, I can only conclude that it wasn't worth anything.
- jumploops 5mo agoAt the end of the article, the author has this to say: > UPDATE: Launch was a success! 400K+ views, and multiple companies reached to use my IP. Read more here[0] [0]https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distribution-fine-tuning/ https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distri...
- komali2 5mo agoPost-hoc justification. There's no analysis of whether that level of hardware was necessary to launch, only that they did get that hardware and did launch.
- jumploops 5mo agoLooking at the GPU utilization graph, it certainly seems like the hardware was saturated for many days/weeks on end. Was it worth it to spend that amount up front, yak shave while building the system, etc. vs. pay for cloud GPUs? Probably not in terms of dollars, when their time is also valued in dollars. Was it worth it for this person? It seems, unequivocally, yes.
- selicos 5mo agoTheir more recent post seems to suggest it was worthwhile. https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distribution-fine-tuning/ https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distri... Abstract/TLDR: LLMs are notoriously formulaic at writing, overusing certain tokens or phrases. I show that models trained with SFT fail to match the distribution of the training data by using Maximum Mean Discrepancy (MMD), Judge Model Quality (JMQ), and L2 Token Distribution.
- kgeist 5mo agoI administer a simple AI server in the office, which just uses a single RTX 5090 but is able to serve ~80 people throughout the day. I'm impressed by Qwen3.6-27b's capabilities in agentic coding/tasks so far. Devs say it's not much different from Sonnet 4.6 on many tasks (sometimes it even outperformed it), 40-60 tok/sec, up to 260k context. The server cost about $10k with all the bells and whistles. I spent a lot of time researching/adding/benchmarking many custom modifications to the software stack and its settings to make the server optimally handle the load with just 1 RTX 5090 without losing quality, but it's still not enough, and the wait times in the queue are getting longer. We're at the limits of the hardware, and I'm out of tricks. The experiment was kind of a success, and the CTO agrees we should scale it. With our own infra, we could run agents 24/7 on everything. Currently, a lot of use cases for the cloud providers are completely blocked by PII/trade secret concerns (our infosec department doesn't buy the "zero retention" promise), plus you don't have to think about billing/budgets/etc. anymore. Now I can't decide how to scale it. On one hand, I'd like to run larger models. And we have the budget to buy, say, 8xH200. But in many benchmarks, the larger models that do fit in 8xH200 comfortably and can serve many parallel requests with acceptable speed/quality don't seem to outperform Qwen3.6 that much in agentic coding/tasks to justify the price. So another option is just to buy a bunch of RTX 6000s and scale horizontally instead: run a copy of a midrange LLM like Qwen3.6 on each GPU. It's cheaper and easier to scale/replace, but then we'll run into problems running larger models in the future if we have to, because of no NVLink support (say, if Alibaba & Co. stop releasing ~30b models and/or ~30b models start falling behind 400b+ models considerably) Does anyone here have experience running large models in a multi-GPU setup with several RTX 6000s in a high-concurrency regime and with large context lengths? (something like Deepseek 4 Flash, Minimax 2.7 etc.)
- zozbot234 5mo agoWouldn't that be a fairly ideal setup for layer parallelism? That doesn't need the high-performance communication of tensor parallelism, and the high-concurrency regime would make it easy to keep the pipeline full with microbatches. You'd also be able to scale out your KV cache storage since that naturally splits layer-wise.
- CobaltFire 5mo ago
- knicholes 5mo agoHe didn't consider the possibility of renting it out during the downtime to Vast.ai to make some money back.
- andix 5mo agoQuick tip for people who want to experiment with local models: A lot of the common smaller models are also available on openrouter or other services. Dirt cheap. I know it's not the same. But a lot of people buy expensive GPUs, just to find out they have no real use for smaller models.
- Cakez0r 5mo agoOpenrouter is great for experimenting with models. I did exactly what you're saying to test smaller models that will run on commodity hardware and determine if it might be worth it to drop $10k on hardware. For me the answer was no, but it's close. I'm very excited for the next few innovation cycles to arrive.
- calmbonsai 5mo agono
- rnxrx 5mo agoThe $48K also isn't fully sunk cost - there's a non-trivial residual value for those GPUs at the moment and likely for a few years yet. The server has a depreciation curve that's pretty enviable, actually!
- jmward01 5mo agoSounds fun/stressful/rewarding. I'm most interested in the update at the end though 'Launch was a success! 400K+ views, and multiple companies reached to use my IP.' I too, like probably 1 in 5 of the people reading this, think I have figured out some major problems with LLMs (context and computation research) but have wondered the best way to 'release' and get value out of it. I can see training being a little easier in that you release weights against a known model arch but not the training code. Wy stuff is all custom layers though. Any thoughts on a release strategy where you need to release the layer code for people to see test weights/the benefits?
- rosmine 5mo agoMy first advice is to have a test set with clear improvement, and a clear "wow" demo use case. There are lots of "breakthroughs" that seem good but aren't (e.g. some new architecture that doesn't mask past tokens correctly and leaks information), so people will assume it is wrong. To prevent this, you need to be extremely rigorous in your launch materials. If you can make it into a product that people can try out themselves, that goes a long way. You don't need to open source any code (I haven't yet) if people can try it out some other way like a demo website. Good luck! Ping me if you want to chat more
- a1o 5mo agoThis is interesting but I am unsure how you make money out of this home setup, I would imagine if one would be offering consultancy to a business the business would make their own equipment/infrastructure available, which would also give a better control of their data. But perhaps I am thinking this because I am thinking about very big companies. Then, on very small business I don’t see they having the use case with the budget to match the need. So is this for specific services for medium sized businesses? Can you explain this a bit?
- dempedempe 5mo agoAt then end they briefly mentioned how they started a service to post-train LLMs on producing more human, less formulaic-obviously-AI text.
- theYipster 5mo agoGreat article. I'm about to embark on a similar journey.... Doing a ton of AI development right now. Don't need a server, but a very, very high end workstation is super appealing to me right now. Looking at $50-$80k. 1TB RAM. 2x RTX Pro 6000s. 64 core Threadripper Pro. As many 4tb or 8tb nvme drives as I can stuff. I envision NixOS at the core... then everything I need virtualized on top with KVM/QEMU. Maybe a dual boot setup with Windows for gaming and Flight Simulator (but I could virtualize that too with easy GPU passthrough.) Lingering questions I'm working to figure out: - Will 2 RTX Pro 6000s run on a 1600 watt PSU? Not sure how much higher I can go without calling an electrician. (standard US home.) - Assuming I plop this into my home office, should I expect the PC to run significantly hotter than my current rig? (3960x threadripper, 128GB RAM, 1600watt psu, overclocked and watercooled 4090.) My water temp, measured at radiator, is about 60c at peak load. (This is the only number I care about, as this is what I have to consider to be comfortable sitting next to it.)
- arjie 5mo agoWhat do you want to do with the workstation? I have a similar setup: - 512 GB - Epyc 9684x - 2x RTX 6000 Pro - 1400 W PSU x 2 but in redundant mode Mine is in a colo where it stays nice and cool. In my case, I went with less RAM and more GPUs (bought 4). Secondarily, the Max-Q blower version of an RTX 6000 Pro Blackwell is easier to keep cool and also only needs 300 W at the cost of very little performance. The non-max-q also only really use 300 W during inference, but the good thing about a lower power use is you can put more GPUs in very safely. I assume you want the Threadripper Pro to maximize single-core performance? So you're spending a lot of time on CPU? Interesting stuff. I gained a lot putting the machine somewhere else. TTFT on a thing like this is between 100-800 ms depending on batching and model size and so on, and your nearest datacenter is likely <10 ms. It sits on nice dual redundant power in a place where it's blown icy cool. Good luck with your setup. If you get around to it, and end up writing about your setup on a blog, do share. Email in profile.
- theYipster 5mo agoVery nice. Primary use case is application development, where the applications leverage a mixture of cloud based and local models. Modelling complex architectures. My work is primarily in the aerospace and defense arena, so hybrid and on-prem are important, as are ITAR and CMMC compliance. The idea is to have the local rig to build and validate architectural deployments that can sit on prem on customer hardware, in cloud, in gov cloud, or in a mix. Not really looking at colocation, as this machine would double as a heavy duty gaming and flight sim rig. That means at least one regular RTX 6000 Pro. Not sure if I can mix and match with the Max-Q version, or if I even want blower fans in a desktop case (last time I did that was about 16-18 years ago with an ATI card... wasn't a fan--pun intended.)
- senectus1 5mo agoYou guys are nuts... I hope you're making enough money to justify this level of investment and power use (not to mention noise and heat management) in your home... I'm just putting a 2nd hand 12gb 3060 into my lab box, but its only for use with HA/Paperless/Plex etc type things. I dont need multi-model agentic behavior for private use. If I did I reckon I'd renting infrastructure rather than filling my home with that sort of gear.
- segmondy 5mo agoThere's a reason folks like these can afford this and you can't. Go on with your cheap self
- sfourdrinier 5mo agoI'll buy it from you!
- zenai666 5mo ago[flagged]
- HWR_14 5mo agoThe other advantage of the local GPU is that you are not feeding your data into cloud providers. I'm not sure how much you can really trust Anthropic and OpenAI not be improving their models based on your input.
- NiloCK 5mo agoDoesn't it benefit me if the models I use improve?
- SJMG 5mo agoDo you value the infinitesimal improvement of model quality more than your privacy?
- deleted 5mo ago[deleted]
- agentifysh 5mo agoYou can turn off training with codex an gemini. Not sure about Anthropic.
- HWR_14 5mo agoI can check a box. How much do you trust OpenAi or Anthropic to not use it as training data anyway? What if you are building a startup and they can just use their visibility into it to copy your IP instead of buying your company?
- flowersjeff 5mo ago"The point of buying the server wasn’t to save money, it was to build something cool." In the end, this is always the real answer - one that I'm sure we can all agree is the 'correct' one too.
- nakedneuron 5mo agoI'm sure thr plan was to build a holodeck all the time..
- am17an 5mo agoPeople doing economics with the cloud GPUs, of course cloud GPUs are going cheaper. But also, is generating tokens all you do with your computer? I can play games on DGX spark and also do LLM inference, so sometimes the economics work out, apart from having fun with it.
- deleted 5mo ago[deleted]
- doh 5mo agoI built a very similar server myself [0] with a similar setup. I run different models for different purposes, but the primary one currently is kimi 2.6. I run kimi as the orchestrator model and then qwen, Gemma and others for specific tasks (sometimes loaded dynamically based on the task at hand), all exposed through the pi harness. I also use Hermes for some personal repeated tasks which connects to the same models, hosted on my local Mac Studio. I am not even going to pretend that this is financially reasonable option. I simply wanted to have a local models. Maybe down the line, as cloud models become less subsidized, I might benefit from having a local setup, but for now, it wasn't the most prudent financial decision. But one big benefit is that I never have worry about my account being randomly banned nor I have to worry about running out of quota. I still use codex and opus for some specific tasks, but as tools are improving, I need them less and less. [0] https://x.com/synopsi/status/2024235558193811778?s=20 https://x.com/synopsi/status/2024235558193811778?s=20
- jwilliams 5mo ago> The mentality shift of renting vs. owning the gpus is huge. When renting, each experiment costs money and I had to ask myself is it worth it. When owning, it feels like not running experiments is costing me money. I feel like there is some very deep generalizable wisdom buried here.
- thomasmarcelis 5mo agoAlso something about subscriptions vs pay-for-usage. I feel the need to use all my weekly tokens or I'm wasting and I bet they would never get this kind of usage out of me if AI ended up being same price per token.
- dgellow 5mo agoI always buy software/assets/dev tools for my hobbies (like CAD, music production, game dev) instead of paying subscriptions, even if that would very likely be way, way cheaper and would give me access to really cool tools. I don’t want to feel bad not using something and I know that’s the case with a subscription
- conormcleod 5mo agoYou should be able to achieve this mentality shift without owning a GPU. You just need to commit some money upfront to cloud GPU spend, in a way that is not feasible to go back on. That way you get the experimentation-encouraging mentality shift to "If I don't use this, I'm wasting my money", without the cost inefficiencies associated with actually buying an accelerator, discussed by others in this thread -> you'll never be able to match the the utilisation and thus the cost amortisation of cloud GPUs.
- zekrioca 5mo agoCan anyone recommend what types of server would it be required to run a RTX 6000 Pro?
- _alphageek 5mo agoI ve seen already one question like that in the thread. But I rephrase it slightly sharper. Did you consider renting out you setup to vast.ai and if so, how much money it can generate per month deducting electricity. Also, sorry for the noob question, is not such server generate enormous amount of heat? You did not use any special cooling system?
- ElenaDaibunny 5mo agothat adds up fast over a year.
- kriro 5mo agoNice analysis, I would have loved a short overview of the kinds of experiments that were running on the machine (I know the results are given). I find the "independent researcher" business model quite interesting. In the linked post he writes """DFT is a proprietary training algorithm, however, I’m currently offering a beta for a model training service where I will train your model for you using DFT.""" I'm curious how successful this is. Essentially market some AI breakthrough as a service instead of publishing a paper like my academic brain is trained to do. As an aside, one thing that I always loved about our field was that the startup cost for many business ideas was "a laptop, internet connection and some some grit". In the age of AI it's quite a bit more and I feel one of the sad side effects of this is that it crowds out poorer and younger developers.
- bob1029 5mo agoAny kind of fixed capacity usage model seems to be a dead end. Paying per token might seem like an exploitative arrangement at first glance, but it's a luxury if you are experimenting or deploying greenfield. Provisioned capacity is a really high end thing. I feel like you'd need to be spending more than $1000/day on tokens for this model to make any sense. You lose a lot of flexibility once you start dumping capital into specific pieces of hardware. Maybe start by renting the GPU server for a few days...
- harrouet 5mo agoThese top AI "independent researchers" that live in underpowered apartments and work off their parents' basement... Is that California ?
- KronisLV 5mo agoThat’s very cool and very expensive - I think the cadastre value of the apartment that I live in is like 35k EUR or thereabout. I wonder how much worse just a bunch of Intel Arc B70s might have been, software fuckery aside. Ofc if I’d need to run local inference or simple fine tunes and learning stuff, I’d probably get one of the SFF options - Mac Minis and all of those Sparks or new AMD AI chips. Then again, I’m broke so go figure. I just fork over some money every month to Anthropic, have been trying out more DeepSeek and also Mistral (their Vibe tool is surprisingly passable under WSL).
- askl 5mo agohttps://en.wikipedia.org/wiki/Betteridge's_law_of_headlines https://en.wikipedia.org/wiki/Betteridge's_law_of_headlines
- remix2000 5mo ago"quit my FAANG job" as in they simultaneously worked in Facebook, Amazon, Apple, Nvidia and Google? Or did the op work at Netflix and is too ashamed to admit that :P
- Frusn 5mo ago[flagged]
- zkmon 5mo agoJust curious - What exactly are you using that rig for? I see that you said research work. Are you building a product or training models? I ask because whether something is worth it or not depends largely on what you get out oof it and how you value what you get. It's perfectly fine to leave a FANG job and go for, say, pottery hobby. What gives you happiness and your value system - these will qualify your decisions.
- lucrbvi 5mo agoIn the article the author says they are doing reinforcement learning with LLMs.
- rob 5mo agoSeems like they just want to play PewDiePie after making tons of money from their salaried job and have a bunch of spare time now.
- titanomachy 5mo agoThey posted their research results, it's linked at the end of the article. https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distribution-fine-tuning/ https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distri...
- bulgur999 5mo ago“Being less able to detect whether a text is AI-generated is exactly what nobody asked for — except villains.” ("—")
- craftedcode 5mo ago[flagged]
- nicman23 5mo agowell as i need to process medical documents i really should not use anything offsite. privacy has a steep cost
- pshirshov 5mo agoQuestions: 1) Was the energy bill factored in? 2) Have you extracted any comparable value out of this?
- DeathArrow 5mo agoThat's a nice problem to have. I can't afford a $48K GPU server, even if I worked as a developer since 25 years ago, because I live in the wrong place.
- bcjdjsndon 5mo ago> if more powerful GPUs could help me make my work be successful just 2 months earlier than I would have with a smaller machine, then buying a more powerful server would be worth it. Jesus I got a migraine trying to parse that
- Radle 5mo ago"If I were to do this again, I wouldn’t do a custom build like this. I would buy a standard datacenter server and rent space in a colocation center. But then I would miss saying Hi to grumbl once in a while."
- reedf1 5mo agoBuy one of these next time, https://tinygrad.org/#tinybox https://tinygrad.org/#tinybox. At least geohot knows what he is doing.
- TeriyakiBomb 5mo agoIt reminds me of the Crypto mining bubble - I look forward to buying my heavily discounted Mac Studio soon.
- b65e8bee43c2ed0 5mo agoa P40 from 10 years ago costs 5x what it did in early 2023. a 3090 from 6 years ago costs as much now as it did then. RAM costs over three times what it did a year ago. M3 512gb now costs twice what it did at release. "soon" isn't happening until China saves us from the Taiwanese semiconductor cabal and their Talmudic markups.
- z3t4 5mo agoCould turn the system into a multi-seat gaming rig with 6 separate gaming seats, using loginctl on Linux.
- cold_harbor 5mo agomissing from most of these cost discussions: privacy. for some workloads the entire value of local is zero data leaving the network, and cloud cost is irrelevant
- mujib77 5mo agoLove to see what u gonna build on $48k GPU
- indianbunghole 5mo ago[dead]
- krupan 5mo agoAnd the net result is a way for LLMs to use more variety in their writing style. Didn't Sam Altman create LLMs to cure cancer and stuff? Why does their writing style matter as long as the information they are conveying is accurate?
- samhoss93 5mo ago[dead]
- adrian_b 5mo agoThe fine print at the end, with "Advice/Other notes", is the most interesting.
- atlasagentsuite 5mo ago[dead]
- teiferer 5mo agoWorked a FAANG job, had the money to burn on a $48k rig. Moves in with parents so they carry the cost of running that rig in their basement. "Totally worth it"
- muralisid 5mo ago[flagged]
- tisiwit562 5mo agodf
- johalmed 4mo ago[flagged]