14 ms·
Don't rent the cloud, own instead
- monster_truck 8mo agoDon't even have to go this far. Colocating in a couple regions will give you most of the logistical thrills at a fraction of the cost!
- coffeebeqn 8mo agoHeavy ML workloads make this more worthwhile since you get to design it to squeeze value out of every facet. For a basic web server and database it’s definitely overkill and something like a colocation makes much more sense
- sys42590 8mo agoIt would be interesting to hear their contingency plan for any kind of disaster (most commonly a fire) that hits their data center.
- sschueller 8mo agoYep, does anyone remember the OVH fire[1][2]? [1] https://www.techradar.com/news/remember-the-ovhcloud-data-center-fire-heres-why-it-was-so-bad https://www.techradar.com/news/remember-the-ovhcloud-data-ce... [2] https://blocksandfiles.com/wp-content/uploads/2023/03/ovhcloud_fire_image_from_bea-ri_report.jpg https://blocksandfiles.com/wp-content/uploads/2023/03/ovhclo...
- AndroTux 8mo agocontingency plan: Don't build your data center out of wood.
- srg0 8mo agoPlastic is made from the same stuff as gasoline.
- direwolf20 8mo agoDrain cleaner and hydrochloric acid makes salt water. Water is made of highly explosive hydrogen. Salt is made of toxic chlorine and explosive sodium.
- otherme123 8mo agoI fully lost three small VPS there, and their response was poor: they didn't even refund time lost, they didn't compensate for time lost (e.g. a couple of months of free VPS), I got better updates from the news than from them (news were saying "almost total loss", while them were trying to convince me that I had the incredible bad luck that my three VPS were in the very small zone affected by the fire). The only way I had to recover what I lost was backups in local machines. When someone point out how safe are cloud providers, as if they have multiple levels of redundancy and are fully protected against even an alien invasion, I remember the OVH fire.
- wiether 8mo agoOVH VPS is not the same as say, AWS EC2. It's their "Compute" under "Public Cloud" that is competing against AWS EC2. https://us.ovhcloud.com/public-cloud/compute/ https://us.ovhcloud.com/public-cloud/compute/ They handled the fire terribly and after that they improved a bit, but an OVH VPS is just a VM running on a single piece of hardware. Quite not the same thing as the "Compute" which is running on clusters.
- instagib 8mo agoFlooding due to burst frozen pipe, false sprinkler trigger, or many others. Something very similar happened at work. Water valve monitoring wasn’t up yet. Fire didn’t respond because reasons. Huge amount of water flooded over a 3 day weekend. Total loss.
- twelvechairs 8mo agoTheres only one solution to this problem and its 2 data centres in some way or form
- mbreese 8mo agoWhat's the line from Contact? why build one when you can have two at twice the price? But, if you're building a datacenter for $5M, spending $10-15M for redundant datacenters (even with extra networking costs), would still be cheaper than their estimated $25M cloud costs.
- golem14 8mo agoOr build two 2.5MM DCs (if can parallelize your workload well enough) and in case of disaster, you only lose capacity. You need however plan for 1MM+ pa in OPEX because good SREs ain’t cheap (or hardware guys building and maintaining machines)
- fpoling 8mo agoThey use the datasenter for model training, not to serve online users. Presumably even if it will be offline for a week or even a month it will not be a total disaster as long as they have, for example, offsite tape backups.
- direwolf20 8mo agothe plan is to not set it on fire. If your office burns down you are already screwed
- langarus 8mo agoThis is a great solution for a very specific type of team but I think most companies with consistent GPU workloads will still just rent dedicated servers and call it a day.
- hyperbovine 8mo agoI agree, and cloud compute is poised to become even more commoditized in the coming years (gazillion new data centers + AI plateauing + efficiency gains, the writing is on the wall). There’s no way this makes sense for most companies.
- NitpickLawyer 8mo ago> AI plateauing Ummm is that plateauing with us in the room? The advantage of renting vs. owning is that you can always get the latest gen, and that brings you newer capabilities (i.e. fp8, fp4, etc) and cheaper prices for current_gen-1. But betting on something plateauing when all the signs point towards the exact opposite is not one of the bets i'd make.
- lelanthran 8mo ago> Ummm is that plateauing with us in the room? Well, the capabilities have already plateaued as far as I can tell :-/ Over the next few yeas we can probably wring out some performance improvements, maybe some efficiency improvements. A lot of the current AI users right now are businesses trying to on-sell AI (code reviewers/code generators, recipe apps, assistant apps, etc), and there's way too many of them in the supply/demand ratio, so you can expect maybe 90% of these companies to disappear in the next few years, taking the demand for capacity with them.
- ocdtrekkie 8mo agoIt's the opposite. The more consistent your workload the more practical and cost-effective it is to go on-prem. Cloud excels for bursty or unpredictable workloads where quickly scaling up and down can save you money.
- cgsmith 8mo agoI used to colocate a 2U server that I purchased with a local data center. It was a great learning experience for me. Im curious why a company wouldn't colocate their own hardware? Proximity isnt an issue when you can have the datacenter perform physical tasks. Bravo to the comma team regardless. It'll be a great learning experience and make each person on their team better. Ps... bx cable instead of conduit for electrical looks cringe.
- vidarh 8mo agoThe main reason not to colocate is if you're somewhere with high real estate costs... E.g Hetzner managed servers competes on price w/co-location for me because I'm in London.
- doublerabbit 8mo agoI colocate in London, a single server / firewall comes to around £5k a year. I also colocate two other servers in some northern UK location in some industrial estate for £2k as my backups. I've never enjoyed the cloud and dedicated server's have their own caveats too. Budget hosts such as Hetzner/OVH have been known to suddenly pull the plug for no reason. My kit is old, second hand old (Cisco UCS 220 M5, 2xDell somethings) and last night I just discovered I can throw in two NVIDIA T4's and turn it in to a personal LLM. I'm quite excited having my own colocated server with basic LLM abilities. My own hardware with my own data and my own cables. Just need my own IP's now.
- vidarh 8mo ago> Budget hosts such as Hetzner/OVH have been known to suddenly pull the plug for no reason. The same would apply for any number of hosts. Hetzner/OVH are cheap, but as your own numbers show the location price gap is more than sufficient to cover the costs of servers. In fact you can colocate with Hetzner too, and you'd get a similar price gap - the lower cost of real-estate is a large part of the reason why they can be as cheap as they are. Data centre operations is a real estate play - to the point that at least one UK data centre operator is owned by a real estate investment company.
- hbogert 8mo agoDatacenters need cool dry air? <45% No, low isn't good perse. I worked in a datacenter which in winters had less than 40%, ram was failing all over the place. Low humidity causes static electricity.
- mbreese 8mo agoLow is good if you are also adding more humidity back in. If you want to maintain 45-50% (guessing), then you would want <45% environmental humidity so that you can raise it to the level you want. You're right about avoiding static, but you'd still want to try to keep it somewhat consistent. It is much cheaper to use external air for cooling if you can.
- hbogert 8mo agoYeah but the article makes it sound as if lower is better, which it is definitely not. And yeah you need to control humidity, that might mean sometimes lowering, and sometimes increase it by whatever solution you have. Also this is where cutting corner indeed results in lower cost, which was the reason for the OP to begin with. It just means you won't get as good a datacenter as people who are actually tuning this whole day and have decades of experience.
- swiftcoder 8mo agoThe datacenter is in San Diego - a quick Google confirms that external humidity pretty much never drops below 50% there. Things would be different in a colder climate where humidity goes --> 0% in the winter
- CamperBob2 8mo agoLow humidity causes static electricity. RAM that is plugged in and operating isn't subject to external ESD, unless you count lightning strikes. Where are you getting this?
- comrade1234 8mo ago15-years ago or so a spreadsheet was floating around where you could enter server costs, compute power, etc and it would tell you when you would break-even by buying instead of going with AWS. I think it was leaked from Amazon because it was always three-years to break-even even as hardware changed over time.
- Onavo 8mo agoWell, somebody should recreate it. I smell a potential startup idea somewhere. There's a ton of "cloud cost optimizers" software but most involve tweaking AWS knobs and taking a cut of the savings. A startup that could offload non critical service from AWS to colo and traditional bare metal hosting like Hetzner has a strong future. One thing to keep in mind is that the curve for GPU depreciation (in the last 5 years at least) is a little steeper than 3 years. Current estimates is that the capital depreciation cost would plunge dramatically around the third year. For a top tier H100 depreciation kicks in around the 3rd year but they mentioned for the less capable ones like the A100 the depreciation is even worse. https://www.silicondata.com/use-cases/h100-gpu-depreciation/ https://www.silicondata.com/use-cases/h100-gpu-depreciation/ Now this is not factoring cost of labour. Labor at SF wages is dreadfully expensive, now if your data center is right across the border in Tijuana on the other hand..
- TonyStr 8mo agoAzure provides their own "Total Cost of Ownership" calculator for this purpose [0]. Notably, this makes you estimate peripheral costs such as cost of having a server administrator, electricity, etc. [0] - https://azure-int.microsoft.com/en-us/pricing/tco/calculator/ https://azure-int.microsoft.com/en-us/pricing/tco/calculator...
- Symbiote 8mo agoI plugged in our own numbers (60 servers we own in a data centre we rent) and Microsoft thinks this costs us an order of magnitude more than it does. Their "assumption" for hardware purchase prices seems way off compared to what we buy from Dell or HP. It's interesting that the "IT labour" cost they estimate is $140k for DIY, and $120k for Azure. Their saving is 5 times more than what we spend...
- Semaphor 8mo agoIn case anyone from comma.ai reads this: "CTO @ comma.ai" the link at the end is broken, it’s relative instead of absolute.
- croisillon 8mo agono because it's on premise you see? you don't need to access the world wide web, just their server /s
- simianwords 8mo agoThe reason companies don’t go with on premises even if cloud is way more expensive is because of the risk involved in on premises. You can see it quite clearly here that there’s so many steps to take. Now a good company would concentrate risk on their differentiating factor or the specific part they have competitive advantage in. It’s never about “is the expected cost in on premises less than cloud”, it’s about the risk adjusted costs. Once you’ve spread risk not only on your main product but also on your infrastructure, it becomes hard. I would be vary of a smallish company building their own Jira in house in a similar way.
- d1sxeyes 8mo agoIt’s also opex vs capex, which is a battle opex wins most of the time.
- simianwords 8mo agoI think it wins because opex is seen as stable recurring cost and capex is seen as the money you put in your primary differentiation for long term gains.
- d1sxeyes 8mo agoTrue, but for a lot of companies “our servers are on-prem” is not a primary differentiator.
- simianwords 8mo agoi think we are saying the same thing?
- TonyStr 8mo agoCapex may also require you to take out loans
- spacebanana7 8mo ago
- camilajets 8mo ago[dead]
- danpalmer 8mo ago> Cloud companies generally make onboarding very easy, and offboarding very difficult. I reckon most on-prem deployments have significantly worse offboarding than the cloud providers. As a cloud provider you can win business by having something for offboarding, but internally you'd never get buy-in to spend on a backup plan if you decide to move to the cloud.
- lelanthran 8mo ago> As a cloud provider you can win business by having something for offboarding, but internally you'd never get buy-in to spend on a backup plan if you decide to move to the cloud. Its the other way around. How do you think all businesses moved to the cloud in the first place?
- danpalmer 8mo agoMy point is that at the point of moving, or creating a new deployment, it's perfectly reasonable to say "how do we get off the cloud if it goes badly", yet no one says "how do we get onto a cloud if managing a datacenter sucks". The cloud providers win business with at least some hint of offboarding support, but on-prem doesn't have that same incentive.
- intalentive 8mo agoI like Hotz’s style: simply and straightforwardly attempting the difficult and complex. I always get the impression: “You don’t need to be too fancy or clever. You don’t need permission or credentials. You just need to go out and do the thing. What are you waiting for?”
- tirant 8mo agoThis was written by Harald Schäfer, the CTO of comma.ai. I'm not so sure if G. Hotz is still involved in comma.ai.
- intalentive 8mo agoAh I missed that.
- piker 8mo agoDon't think he is, but it does seem like he inspired a hacker mentality in the shop during his tenure.
- gogasca 8mo ago[dead]
- jillesvangurp 8mo agoAt scale (like comma.ai), it's probably cheaper. But until then it's a long term cost optimization with really high upfront capital expenditure and risk. Which means it doesn't make much sense for the majority of startup companies until they become late stage and their hosting cost actually becomes a big cost burden. There are in between solutions. Renting bare metal instead of renting virtual machines can be quite nice. I've done that via Hetzner some years ago. You pay just about the same but you get a lot more performance for the same money. This is great if you actually need that performance. People obsess about hardware but there's also the software side to consider. For smaller companies, operations/devops people are usually more expensive than the resources they manage. The cost to optimize is that cost. The hosting cost usually is a rounding error on the staffing cost. And on top of that the amount of responsibilities increases as soon as you own the hardware. You need to service it, monitor it, replace it when it fails, make sure those fans don't get jammed by dust puppies, deal with outages when they happen, etc. All the stuff that you pay cloud providers to do for you now becomes your problem. And it has a non zero cost. The right mindset for hosting cost is to think of it in FTEs (full time employee cost for a year). If it's below 1 (most startups until they are well into scale up territory), you are doing great. Most of the optimizations you are going to get are going to cost you in actual FTEs spent doing that work. 1 FTE pays for quite a bit of hosting. Think 10K per month in AWS cost. A good ops person/developer is more expensive than that. My company runs at about 1K per month (GCP and misc managed services). It would be the wrong thing to optimize for us. It's not worth spending any amount of time on for me. I literally have more valuable things to do. This flips when you start getting into the multiple FTEs per month in cost for just the hosting. At that point you probably have additional cost measured in 5-10 FTE in staffing anyway to babysit all of that. So now you can talk about trading off some hosting FTEs for modest amount of extra staffing FTEs and make net gains.
- durakot 8mo agoThere's the HN I know and love
- kavalg 8mo agoThis was one of the coolest job ads that I've ever read :). Congrats for what you have done with your infrastructure, team and product!
- HanClinto 8mo agoAgreed! Gives a whole new level to the idea of "full stack developer"
- tirant 8mo agoWell, their comment section is fore sure not running on premises, but on the cloud: "An error occurred: API rate limit already exceeded for installation ID 73591946."
- clarity_hacker 8mo ago[dead]
- pja 8mo agoI’m impressed that San Diego electrical power manages to be even more expensive than in the UK. That takes some doing.
- satvikpendem 8mo agoI just read about Railway doing something similar, sadly their prices are still high compared to other bare metal providers and even VPS such as Hetzner with Dokploy, very similar feature set yet for the same 5 dollars you get way more CPU, storage and RAM. https://blog.railway.com/p/launch-week-02-welcome https://blog.railway.com/p/launch-week-02-welcome
- dist-epoch 8mo agoTheir pricing page is so confusing: CPU: $0.00000772 per vCPU / sec This seems to imply $40 / month for 2 vCPU which seems very high? Or maybe they mean "used" CPU versus idle?
- Neil44 8mo agoBilling per used or not idle cpu cycle would be quite interesting. Number of cores would just effectively be your cost cap. Efficiency would be even more important. And if the provider over subscribes cores you just pay less. Actually that's probably why they don't do it...
- efreak 8mo agoDon't most big clouds not share cores between tenants? I have a vague feeling that around spectre/meltdown this was stopped. I wouldn't be surprised to be wrong, but if you're dedicating a core to a VM, you're not going to charge less for unused CPU that nobody else can use.
- speedgoose 8mo agoI would suggest to use both on-premise hardware and cloud computing. Which is probably what comma is doing. For critical infrastructure, I would rather pay a competent cloud provider than being responsible for reliability issues. Maintaining one server room in the headquarters is something, but two servers rooms in different locations, with resilient power and network is a bit too much effort IMHO. For running many slurm jobs on good servers, cloud computing is very expensive and you sometimes save money in a matter of months. And who cares if the server room is a total loss after a while, worst case you write some more YAML and Terraform and deploy a temporary replacement in the cloud. Another thing between is colocation, where you put hardware you own in a managed data center. It’s a bit old fashioned, but it may make sense in some cases. I can also mention that research HPCs may be worth considering. In research, we have some of the world fastest computers at a fraction of the cost of cloud computing. It’s great as long as you don’t mind not being root and having to use slurm. I don’t know in USA, but in Norway you can run your private company slurm AI workloads on research HPCs, though you will pay quite a bit more than universities and research institutions. But you can also have research projects together with universities or research institutions, and everyone will be happy if your business benefits a lot from the collaboration.
- olavgg 8mo ago> I would rather pay a competent cloud provider than being responsible for reliability issues. Why do so many developers and sysadmins think they're not competent for hosting services. It is a lot easier than you think, and its also fun to solve technical issues you may have.
- pageandrew 8mo agoThe point was about redundancy / geo spread / HA. It’s significantly more difficult to operate two physical sites than one. You can only be in one place at a time. If you want true reliability, you need redundant physical locations, power, networking. That’s extremely easy to achieve on cloud providers.
- account42 8mo ago
- rvz 8mo agoNot long ago Railway moved from GCP to their own infrastructure since it was very expensive for them. [0] Some go for a Oxide rack [1] for a full stack solution (both hardware and software) for intense GPU workloads, instead of building it themselves. It's very expensive and only makes sense if you really need infrastructure sovereignty. It makes more sense if you're profitable in the tens of millions after raising hundreds of millions. It also makes sense for governments (including those in the EU) which should think about this and have the compute in house and disconnected from the internet if they are serious about infrastructure sovereignty, rather than depending on US-based providers such as AWS. [0] https://blog.railway.com/p/data-center-build-part-one https://blog.railway.com/p/data-center-build-part-one [1] https://oxide.computer/ https://oxide.computer/
- rasjani 8mo agoI was under impression that Oxide rack does not currently ship with GPU's - at least with buildin. . Has this changed recently ?
- panick21_ 8mo agoOxide racks don't yet have a GPU solution. But it is a good options for general compute and even with GPU required, general compute hasn't gone away.
- kaon_2 8mo agoAm I the only one that is simply scared of running your own cloud? What happens if your administrator credentials get leaked? At least with Azure I can phone microsoft and initiate a recovery. Because of backups and soft deletion policies quite a lot is possible. I guess you can build in these failsafe scenarios locally too? But what if a fire happens like in South Korea? Sure most companies run more immediate risks such as going bankrupt, but at least Cloud relieves me from the stuff of nightmares. Except now I have nightmares that the USA will enforce the patriot act and force Microsoft to hand over all their data in European data centers and then we have to migrate everything to a local cloud provider. Argh...
- vachina 8mo agoThen literally own the cloud, like run the hardware on-prem yourself.
- direwolf20 8mo agoDo you have a computer at home? Are you scared of its credentials leaking? A server is just another computer with a good internet connection. You can equip your server with a mouse, keyboard and screen and then it doesn't even need credentials. The credential is your physical access to the mouse and keyboard.
- geodel 8mo agoI mean people are nowadays are really scared of using microwave oven too. What happens if I heat my coffee 1 min too long. Could be near death experience. Thats why I always drive down to Starbucks for coffee!
- direwolf20 8mo agoTrue! Decline of defiance or something. Everyone is suddenly a follower. Any idea what caused it? Micro plastics in the brain? Social media?
- pu_pe 8mo ago> Self-reliance is great, but there are other benefits to running your own compute. It inspires good engineering. It's easy to inspire people when you have great engineers in the first place. That's a given at a place like comma.ai, but there are many companies out there where administering a datacenter is far beyond their core competencies. I feel like skilled engineers have a hard time understanding the trade-offs from cloud companies. The same way that comma.ai employees likely don't have an in-house canteen, it can make sense to focus on what you are good at and outsource the rest.
- szszrk 8mo ago> I feel like skilled engineers have a hard time understanding the trade-offs from cloud companies. They spend too much time on yet another cloud native support group call, learning for ThatOneCloudProvider certificates, figuring out that single implementation caveats, standardizing security procedures between cloud teams, and so on. Yet people in the article just throw a 1000 lines of code KV store mkv [0] on a huge raw storage server and call it a day. And it's a legit choice, they did actual study beforehand and concluded: we don't need redundancy in most cases. At all. I respect that. [0] https://github.com/geohot/minikeyvalue https://github.com/geohot/minikeyvalue
- Torq_boi 8mo agoWe actually do have an in-house chef lol.
- BoredPositron 8mo agocapex vs opex the Opera.
- petesergeant 8mo agoOne thing I don't really understand here is why they're incurring the costs of having this physically in San Diego, rather than further afield with a full-time server tech essentially living on-prem, especially if their power numbers are correct. Is everyone being able to physically show up on site immediately that much better than a 24/7 pair of remote hands + occasional trips for more team members if needed?
- deleted 8mo ago[deleted]
- adamcharnock 8mo agoThis is an industry we're[0] in. Owning is at one end of the spectrum, with cloud at the other, and a broadly couple of options in-between: 1 - Cloud – This is minimising cap-ex, hiring, and risk, while largely maximising operational costs (its expensive) and cost variability (usage based). 2 - Managed Private Cloud - What we do. Still minimal-to-no cap-ex, hiring, risk, and medium-sized operational cost (around 50% cheaper than AWS et al). We rent or colocate bare metal, manage it for you, handle software deployments, deploy only open-source, etc. Only really makes sense above €$5k/month spend. 3 - Rented Bare Metal – Let someone else handle the hardware financing for you. Still minimal cap-ex, but with greater hiring/skilling and risk. Around 90% cheaper than AWS et al (plus time). 4 - Buy and colocate the hardware yourself – Certainly the cheapest option if you have the skills, scale, cap-ex, and if you plan to run the servers for at least 3-5 years. A good provider for option 3 is someone like Hetzner. Their internal ROI on server hardware seems to be around the 3 year mark. After which I assume it is either still running with a client, or goes into their server auction system. Options 3 & 4 generally become more appealing either at scale, or when infrastructure is part of the core business. Option 1 is great for startups who want to spend very little initially, but then grow very quickly. Option 2 is pretty good for SMEs with baseline load, regular-sized business growth, and maybe an overworked DevOps team! [0] https://lithus.eu https://lithus.eu, adam@
- mgaunard 8mo agoyou're missing 5, what they are doing. There is a world of difference between renting some cabinets in an Equinix datacenter and operating your own.
- adamcharnock 8mo agoFair point! 5 - Datacenter (DC) - Like 4, except also take control of the space/power/HVAC/transit/security side of the equation. Makes sense either at scale, or if you have specific needs. Specific needs could be: specific location, reliability (higher or lower than a DC), resilience (conflict planning). There are actually some really interesting use cases here. For example, reliability: If your company is in a physical office, how strong is the need to run your internal systems in a data centre? If you run your servers in your office, then there's no connectivity reliability concerns. If the power goes out, then the power is out to your staff's computers anyway (still get a UPS though). Or perhaps you don't need as high reliability if you're doing only batch workloads? Do you need to pay the premium for redundant network connections and power supplies? If you want your company to still function in the event of some kind of military conflict, do you really want to rely on fibre optic lines between your office and the data center? Do you want to keep all your infrastructure in such a high-value target? I think this is one of the more interesting areas to think about, at least for me!
- juvoly 8mo ago> Cloud companies generally make onboarding very easy, and offboarding very difficult. If you are not vigilant you will sleepwalk into a situation of high cloud costs and no way out. If you want to control your own destiny, you must run your own compute. Cost and lock-in are obvious factors, but "sovereignty" has also become a key factor in the sales cycle, at least in Europe. Handing health data, Juvoly is happy to run AI work loads on premise.
- arjie 8mo agoRealistically, it's the speed with which you can expand and contract. The cloud gives unbounded flexibility - not on the per-request scale or whatever, but on the per-project scale. To try things out with a bunch of EC2s or GCEs is cheap. You have it for a while and then you let it go. I say this as someone with terabytes of RAM in servers, and a cabinet I have in the Bay Area.
- Dormeno 8mo agoThe company I work for used to have a hybrid where 95% was on-prem, but became closer to 90% in the cloud when it became more expensive to do on-prem because of VMware licensing. There are alternatives to VMware, but not officially supported with our hardware configuration, so the switch requires changing all the hardware, which still drives it higher than the cloud. Almost everything we have is cloud agnostic, and for anything that requires resilience, it sits in two different providers. Now the company is looking at doing further cost savings as the buildings rented for running on-prem are sitting mostly unused, but also the prices of buildings have gone up in recent years, notably too, so we're likely to be saving money moving into the cloud. This is likely to make the cloud transition permanent.
- evertheylen 8mo ago> Maintaining a data center is much more about solving real-world challenges. The cloud requires expertise in company-specific APIs and billing systems. A data center requires knowledge of Watts, bits, and FLOPs. I know which one I rather think about. I find this to be applicable on a smaller scale too! I'd rather setup and debug a beefy Linux VPS via SSH than fiddle with various propietary cloud APIs/interfaces. Doesn't go as low-level as Watts, bits and FLOPs but I still consider knowledge about Linux more valuable than knowing which Azure knobs to turn.
- jongjong 8mo agoOr better; write your software such that you can scale to tens of thousands of concurrent users on a single machine. This can really put the savings into perspective.
- swiftcoder 8mo agoIf you were to read TFA, it is about ML training workloads, not web servers
- jongjong 8mo agoWell the article starts out with a suggestion that we should all get a data center... It's quite a jump to assume that everyone reading this article needs to train their own LLMs.
- faust201 8mo agoLook the bottom of that page: An error occurred: API rate limit already exceeded for installation ID 73591946. Error from https://giscus.app/ https://giscus.app/ Fellow says one thing and uses another.
- RT_max 8mo ago[dead]
- yomismoaqui 8mo agoThis quote is gold: The cloud requires expertise in company-specific APIs and billing systems. A data center requires knowledge of Watts, bits, and FLOPs. I know which one I rather think about.
- rudolph9 8mo ago> Having your own data center is cool This company sounds more like a hobby interest than a business focused on solving genuine problems.
- BirAdam 8mo agoTo me it sounds more like a return to vertical integration. This is becoming increasingly common as far as I can tell. There are benefits either direction, and I think that each company needs to evaluate the pros and cons themselves. Emotional pros/cons are something companies need to evaluate as employee morale can make or break a company. If the company is super technical in culture and they gain something intangible that is boosting the bottom line, having a datacenter as a "cool" factor is probably worth it.
- vovavili 8mo agoI'd argue that it is in the long-term interest of any genuinely innovative company to attract intellectually curious talent with some coolness factor.
- HanClinto 8mo agoIt kinda' does, doesn't it? Re: the "hobby" part is where I agree with you the most. Where you say it's not solving genuine problems is where I differ the most. It really feels to me like Comma is staffed by people who recognize that they never stopped enjoying playing with Lego -- their bricks just grew up, and they realized they can: 1) solve real-world problems 2) not be jerks about it 3) get paid to do it Not everything has to be about optimizing for #3. I'm a happy paying customer of Comma.ai (Comma four, baby!) -- their product is awesome, extremely consumer-friendly, and I hope they can grow in their success!
- dagi3d 8mo ago> San Diego power cost is over 40c/kWh, ~3x the global average. It’s a ripoff, and overpriced simply due to political dysfunction. Mind anyone elaborate? Always thought this is was a direct cause of the free market. Not sure if by dysfunction the op means lack of intervention.
- amluto 8mo agoDid you say “free market”? There is one provider. There is a lot of regulation, mostly incompetent. It’s a mess.
- throwawaypath 8mo ago>Mind anyone elaborate? Always thought this is was a direct cause of the free market. Not sure if by dysfunction the op means lack of intervention. The majority of Californians have no say and cannot choose their utilities provider. This is the polar opposite of the "free market".
- omoikane 8mo agoElectricity cost in California is generally more expensive than most other US states, except Hawaii. Not sure why. Perhaps Comma needed the datacenter to be in San Diego for latency or other reasons, but if they need it mostly for compute, it would have been cheaper to operate their datacenter elsewhere... but if we keep going down that path, maybe it actually becomes cheaper to rent a cloud after all.
- Maro 8mo agoWorking at a non-tech regional bigco, where ofc cloud is the default, I see everyday how AWS costs get out of hand, it's a constant struggle just to keep costs flat. In our case, the reality is that NONE of our services require scalability, and the main upside of high uptime is nice primarily for my blood pressure.. we only really need uptime during business hours, nobody cares what happens at night when everybody is sleeping. On the other hand, there's significant vendor lockin, complexity, etc. And I'm not really sure we actually end up with less people over time, headcount always expands over time, and there's always cool new projects like monitoring, observability, AI, etc. My feeling is, if we rented 20-30 chunky machines and ran Linux on them, with k8s, we'd be 80% there. For specific things I'd still use AWS, like infinite S3 storage, or RDS instances for super-important data. If I were to do a startup, I would almost certainly not base it off AWS (or other cloud), I'd do what I write above: run chunky servers on OVH (initially just 1-2), and use specific AWS services like S3 and RDS. A bit unrelated to the above, but I'd also try to keep away from expensive SaaS like Jira, Slack, etc. I'd use the best self-hosted open source version, and be done with it. I'd try Gitea for git hosting, Mattermost for team chat, etc. And actually, given the geo-political situation as an EU citizen, maybe I wouldn't even put my data on AWS at all and self-host that as well...
- vasco 8mo agoHaving worked only with the cloud I really wonder if these companies don't use other software with subscriptions. Even though AWS is "expensive" its a just another line item compared to most companies overall SaaS spend. Most businesses don't need that much compute or data transfer in the grand scheme of things.
- bob1029 8mo agoThe #1 reason I would advocate for using AWS today is the compliance package they bring to the party. No other cloud provider has anything remotely like Artifact. I can pull Amazon's PCI-DSS compliance documentation using an API call. If you have a heavily regulated business (or work with customers who do), AWS is hard to beat. If you don't have any kind of serious compliance requirement, using Amazon is probably not ideal. I would say that Azure AD is ok too if you have to do Microsoft stuff, but I'd never host an actual VM on that cloud. Compliance and "Microsoft stuff" covers a lot of real world businesses. Going on prem should only be done if it's actually going to make your life easier. If you have to replicate all of Azure AD or Route53, it might be better to just use the cloud offerings.
- wiether 8mo ago> The #1 reason I would advocate for using AWS today is the compliance package they bring to the party. I was going to post the same comment. Most of the people agreeing to foot the AWS bill do it because they see how much the compliance is worth to them.
- mrbluecoat 8mo agoStopped reading at "Our main storage arrays have no redundancy". This isn't a data center, it's a volatile AI memory bank.
- huntaub 8mo agoThis turns out to be a more and more important primitive for companies who are building their own models [1]. [1] https://si.inc/posts/the-heap/ https://si.inc/posts/the-heap/
- sgarland 8mo agoYou should have kept reading: > Redundancy is not needed since no specific data is critical. > we have a redundant mkv storage array to store all of our trained models and training metrics. That's just called understanding your failure domains, and RTO/RPO needs.
- architsingh15 8mo agoLooks insanely daunting imo
- deleted 8mo ago[deleted]
- CodeCompost 8mo agoMicrosoft made the TCO argument and won. Self-hosting is only an option if you can afford expensive SysOps/DevOps/WhateverWeAreCalledTheseDays to manage it.
- davsti4 8mo agoSo.... you're saying they must be understaffed and paying poverty range wages to afford the San Diego climate and still cut a profit? ;)
- macmac_mac 8mo agoChatgpt: # don’t own the cloud, rent instead the “build your own datacenter” story is fun (and comma’s setup is undeniably cool), but for most companies it’s a seductive trap: you’ll spend your rarest resource (engineer attention) on watts, humidity, failed disks, supply chains, and “why is this rack hot,” instead of on the product. comma can justify it because their workload is huge and steady, they’re willing to run non-redundant storage, and they’ve built custom GPU boxes and infra around a very specific ML pipeline. ([comma.ai blog][1]) ## 1) capex is a tax on flexibility a datacenter turns “compute” into a big up-front bet: hardware choices, networking choices, facility choices, and a depreciation schedule that does not care about your roadmap. cloud flips that: you pay for what you use, you can experiment cheaply, and you can stop spending the minute a strategy changes. the best feature of renting is that quitting is easy. ## 2) scaling isn’t a vibe, it’s a deadline real businesses don’t scale smoothly. they spike. they get surprise customers. they do one insane training run. they run a migration. owning means you either overbuild “just in case” (idle metal), or you underbuild and miss the moment. renting means you can burst, use spot/preemptible for the ugly parts, and keep steady stuff on reserved/committed discounts. ## 3) reliability is more than “it’s up most days” comma explicitly says they keep things simple and don’t need redundancy for ~99% uptime at their scale. ([comma.ai blog][1]) that’s a perfectly valid trade—if your business can tolerate it. many can’t. cloud providers sell multi-zone, multi-region, managed backups, managed databases, and boring compliance checklists because “five nines” isn’t achieved by a couple heroic engineers and a PID loop. ## 4) the hidden cost isn’t power, it’s people comma spent ~$540k on power in 2025 and runs up to ~450kW, plus all the cooling and facility work. ([comma.ai blog][1]) but the larger, sneakier bill is: on-call load, hiring niche operators, hardware failures, spare parts, procurement, security, audits, vendor management, and the opportunity cost of your best engineers becoming part-time building managers. cloud is expensive, yes—because it bundles labor, expertise, and economies of scale you don’t have. ## 5) “vendor lock-in” is real, but self-lock-in is worse cloud lock-in is usually optional: you choose proprietary managed services because they’re convenient. if you’re disciplined, you can keep escape hatches: containers, kubernetes, terraform, postgres, object storage abstractions, multi-region backups, and a tested migration plan. owning your datacenter is also lock-in—except the vendor is past you, and the contract is “we can never stop maintaining this.” ## the practical rule *if you have massive, predictable, always-on utilization, and you want to become good at running infrastructure as a core competency, owning can win.* that’s basically comma’s case. ([comma.ai blog][1]) *otherwise, rent.* buy speed, buy optionality, and keep your team focused on the thing only your company can do. if you want, tell me your rough workload shape (steady vs spiky, cpu vs gpu, latency needs, compliance), and i’ll give you a blunt “rent / colo / own” recommendation in 5 lines. [1]: https://blog.comma.ai/datacenter/ https://blog.comma.ai/datacenter/ "Owning a $5M data center - comma.ai blog"
- segmondy 8mo agoI cancelled my digital ocean server of almost a decade late last year and replaced it with a raspberry pi 3 that was doing nothing. We can do it, we should do it.
- Havoc 8mo agoInteresting that they go for no redundancy
- figmert 8mo agoWhat redundancy are we talking about? AWS has proven to the world on multiple occasions that redundancy across geo locations is useless, because if us-east-1 is down, their whole cloud is done, causing a big chunk of the world to be down. Half sarcasm of course, but it goes to show that the world is not going to fall apart in many cases when it comes to software. Sure, it's not ideal in lots of cases, but we'll survive without redundancy.
- bovermyer 8mo agoI'm thinking about doing a research project at my university looking into distributed "data centers" hosted by communities instead of centralized cloud providers. The trick is in how to create mostly self-maintaining deployable/swappable data centers at low cost...
- nubela 8mo agoSame thing. I was previously spending 5-8K on DigitalOcean, supposedly a "budget" cloud. Then the company was sold, and I started a new company on entirely self-hosted hardware. Cloudflare tunnel + CC + microk8s made it trivial! And I spend close to nothing other than internet that I already am spending on. I do have solar power too.
- nickorlow 8mo agoEven at the personal blog level, I'd argue it's worth it to run your own server (even if it's just an old PC in a closet). Gets you on the path to running a home lab.
- drnick1 8mo agoAbsolutely. I don't have a blog but run my own email, several game servers, Matrix instance, Nextcloud and other internal services on a retired gaming PC. The total cost of my cloud subscriptions is $0, and no one is snooping on me. It's a great setup when combined with Linux machines and GrapheneOS phones, completely private and free of Big Tech.
- nottorp 8mo ago> We use SSDs for reliability and speed. Hey, how do SSDs fail lately? Do they ... vanish off the bus still? Or do they go into read only mode?
- MORPHOICES 8mo ago[dead]
- imcritic 8mo agoI love articles like this and companies with this kind of openness. Mad respect to them for this article and for sharing software solutions!
- apothegm 8mo agoThis also depends so much on your scaling needs. If you need 3 mid-sized ECS/EC2 instances, a load balancer, and a database with backups, renting those from AWS isn’t going to be significantly more expensive for a decent-sized company than hiring someone to manage a cluster for you and dealing with all the overhead of keeping it maintained and secure. If you’re at the scale of hundreds of instances, that math changes significantly. And a lot of it depends on what type of business you have and what percent of your budget hosting accounts for.
- infecto 8mo agoI also thinks it’s risk model too. Every time I see these kind of posts I think it misses the point there is a balance not only on cost like you describe but risk as well. You are paying to offload some of the risk from yourself.
- eldenring 8mo agoThe issue is that they have already paid off their datacenter 5x over compared to cloud. For offline, batch training, I don't ses how any amount of risk could offset the savings.
- infecto 8mo agoIt’s no issue and it’s right in the front for their situation, if your business is computer it makes little sense for cloud. That said from the risk perspective I assume for what their doing in the data center there is low risk if downtime happens.
- apothegm 8mo agoModel training, perhaps. Web hosting? Not so much.
- betaby 8mo ago> You are paying to offload some of the risk from yourself. The opposite is also true: one is risking being banned by exascalers.
- lovegrenoble 8mo agoI've just shifted to Hetzner, no regret
- ghc 8mo agoIf it were me, instead of writing all these bespoke services to replicate cloud functionality, I'd just buy oxide.computer systems.
- tucnak 8mo agoOxide doesn't support GPU deployments, or any other accelerators, like SmartNIC or DPU for that matter.
- butterisgood 8mo agoI think this is how IBM is making tons of money on mainframes. A lot of what people are doing with cloud can be done on premises with the right levels of virtualization. https://intellectia.ai/news/stock/ibm-mainframe-business-achieves-best-revenue-in-20-years?utm_source=chatgpt.com https://intellectia.ai/news/stock/ibm-mainframe-business-ach... 60% YoY growth is pretty excellent for an "outdated" technology.
- IFC_LLC 8mo agoThis is cool. Yet, there are levels of insanity and those depend on your inability to estimate things. When I'm launching a project it's easier for me to rent $250 worth of compute from AWS. When the project consumes $30k a month, it's easier for me to rent a colocation. My point is that a good engineer should know how to calculate all the ups and downs here to propose a sound plan to the management. That's the winning thing.
- piker 8mo agoIt goes further than this first order, though. If you're trying to build a business that attracts the types of talent who wants to know the stack up and down, starting with an AWS instance might give you a better shot at funding (and thus a better overall shot), but it's not clear that it gives you a shot a building the business you're aiming for. For the things that "don't make your beer better", sure, but we're talking about training ML models at an ML shop. Here it makes sense for this reason.
- infecto 8mo agoThat last part is exactly it and I while I know the intro sentence nails it I don’t think compute resonates with people (everyone uses compute). If you are 24/7 running work at scale it absolutely makes sense past the initial first couple years to build out your own DC like this.
- redrove 8mo agoWe’re past the point in history where most engineers get to make even a recommendation about which platform to use to management. In 99.999999% of cases management has already decided and is just informing you, because they know better.
- JackSlateur 8mo agoI work in a multi-billions dollars company and do not face what you describe Perhaps an exception (yet so far, I've never encounter the situation you describe)
- squeefers 8mo agomark my words. cloud will fall out of fashion, but it will come back in fashion under another name in some amount of years. its cyclical.
- JKCalhoun 8mo agoNaive comment from a hobbyist with nothing close to $5M: I'm curious about the degree to which you build a "home lab" equivalent. I mean if "scaling" turned out to be just adding another Raspberry Pi to the rack (where is Mr. Geerling when you need him?) I could grow my mini-cloud month by month as spending money allowed. (And it would be fun too.)
- user34283 8mo agoI paid 150€ for a Mini PC with an Intel N100, 16 GB of DDR5 memory, and a 500 GB SSD. While I have no intention to scale up low spec hardware like this, it at least seems to beat the Azure VMs we use at work with "4 CPUs", which corresponds to two physical cores on an AMD EPYC CPU. And that super slow machine I understand costs more than $100 per month, and that's without charges for disk space slower than the SSD, or network traffic. Renting at Azure seems to be a terrible decision, particularly for desktop use.
- sixothree 8mo agoIt's hard to describe how slow a $150 / month azure VM really is. Holy heck are they limiting.
- coffeebeqn 8mo agoYou sure can. Pi are pretty underpowered you can get machines with more cores and memory and pcie lanes and networking out there and virtualize them
- sgarland 8mo agoThe degree is whatever you want to deal with. I had a rack at my last house (need to redesign the space for it at new house) with 3x Dell R620s in a Proxmox cluster, running K8s, serving Ceph from NVMe drives over Infiniband (for the mesh traffic), and 2x Supermicros running independent ZFS pools. It was fun to build - especially Infiniband - but my next iteration is going to be a single beefy server, maybe with storage attached externally. What I had had outstanding uptime, but ultimately it was massively overkill, noisy, hot, and sucked power down.
- infecto 8mo agoI love this article. Great write up. Gave me the same feeling when I would read about Stackoverflows handful of servers that ran all of the sites.
- gwbas1c 8mo agoTLDR: > In comma’s case I estimate we’ve spent ~5M on our data center, and we would have spent 25M+ had we done the same things in the cloud. IMO, that's the biggie. It's enough to justify paying someone to run their datacenter. I wish there was a bit more detail to justify those assumptions, though. That being said, if their needs grow by orders of magnitude, I'd anticipate that they would want to move their servers somewhere with cheaper electricity.
- engelo_b 8mo ago[dead]
- seg_lol 8mo agoUsing "big" cloud providers is often a mistake. You want to use rented assets to bootstrap and then start deploying on instances that are more and more under your control. With big cloud providers, it is easy to just succumb to their service offerings rather than do the right thing. Do your PoC on Hetzner and DigitalOcean then scale with purpose.
- komali2 8mo ago> The cloud requires expertise in company-specific APIs and billing systems. This is one reason I hate dealing with AWS. It feels like a waste of time in some ways. Like learning a fly-by-night javascript library - maybe I'm better off spending that time writing the functionality on my own, to increase my knowledge and familiarity?
- rmoriz 8mo agoCloud, in terms of "other company's infrastructure" always implies losing the competence to select, source and operate hardware. Treating hardware as commodity will eventually treat your very own business as commodity: Someone can just copy your software/IP and ruin your business. Every durable business needs some kind of intellectual property and human skills that are not replaceable easily. This sounds binary, but isn't. You can build long-lasting partnerships. German Mittelstand did that over decades.
- assaddayinh 8mo agoIs there a client to sell on your own unused private cloud?
- bradley13 8mo agoGoes for small business and individuals as well. Sure, there are times that cloud makes sense, but you can and should do a lot on your own hardware.
- kevinkatzke 8mo agoFeels like I’ve lived through a full infrastructure fashion cycle already. I started my career when cloud was the obvious answer and on-prem was “legacy.” Now on-prem is cool again. Makes me wonder whether we’re already setting up the next cycle 10 years from now, when everyone rediscovers why cloud was attractive in the first place and starts saying “on-prem is a bad idea” again.
- devmor 8mo agoSometimes, I feel like this is indicative of the incredible waste present in IT and development. Granted the cost of this kind of infrastructure upheaval is orders of magnitude cheaper than something like manufacturing - but still, it feels ridiculous that established companies can swap back and forth on a whim.
- andrewstuart2 8mo agoThe problem was always the platform. For me, I saw very early on that kubernetes was exactly what I wanted after reading about how Google "treats the datacenter like one large computer." And I've been very happily running my own side projects on my own home cluster for 10 ish years (my kube-system namespace is 9y old). But selling any of my employers on this was a very hard proposition until enough people had shown it working at that scale.
- Aromasin 8mo agoIf this were cyclical, I'd be inclined to agree, but this seems to be more of a wave. I also think the push back is more than just one against rented compute. It is tied to a societal ennui that comes from the feeling that we no longer own anything, be it music, housing, movies, land, tools, phones, or cars. Everything is moving to either being rented or on credit. There's a push back against this self-made feudal revival, and that scales all the way from individuals through to corporations; in this case, against the idea that a mega-corporation gets to decide how and when you get to use your compute, and at what variable price.
- Aurornis 8mo ago> Makes me wonder whether we’re already setting up the next cycle 10 years from now, when everyone rediscovers why cloud was attractive in the first place and starts saying “on-prem is a bad idea” again. My entire career I’ve encountered people passionately pushing for on-prem and railing against anything cloud. I can’t remember a time when Hacker News comments leaned pro-cloud because it’s always been about self-hosting. The few times the on-prem people won out in my career never went exactly as they imagined. Buying a couple servers and setting them up at the colo is easy enough, but the slow and steady drag of maintaining your own infrastructure starts to work its way into every development cycle after that. In my experience, every team has significantly underestimated how all the little things add up to a drag on available time for other work. The best case for on-prem that I saw was when a company was basically in maintenance mode. Engineers had a lot of extra time to optimize, update. maintain, and cost reduce without subtracting from feature development or bug fixes. The worst cases for on-prem I’ve seen have been funded startups. In this situation it’s imperative that everyone focus on feature development and rapid iteration. Letting some of the engineers get sidetracked with setting up and maintaining their own hosting to save a dollar amount that barely hires 1-2 more engineers but sets the schedule back by many months was a huge mistake. In my experience, most engineers become less enchanted with rolling their own on premises hosting as they get older. Their work becomes more about getting the job done quickly and to budget, not hyper-optimizing the hosting situation at the expense of inviting more complexity and miscellaneous tasks into their workload.
- devmor 8mo ago> In a future blog post I hope I can tell you about how we produce our own power and you should too. Rackmounted fusion reactors, I hope. Would solve my homelab wattage issues too.
- drnick1 8mo agoOn premises isn't only about saving money (that's not always clear). The article neglects the most important benefits which are freedom (control) and privacy. It's basically the same considerations that apply to owning vs renting a house.
- Aurornis 8mo agoThe entire second section is about different benefits of having your own data centers. Cost is listed as the last one, not the primary one.
- sgarland 8mo agoNote that they're running R630/R730s for storage. Those are 12-year old servers, and yet they say each one can do 20 Gbps (2.5 GBps) of random reads. In comparison, the same generation of hardware at AWS ({c,m,r}4) instance maxes out at 50% of that for EBS throughput on m4, and 70% on r4 - and that assumes carefully tuned block sizes. Old hardware is _plenty_ powerful for a lot of tasks today.
- treesknees 8mo agoI’m on a project at work replacing our R430s and R730s. They’ve been absolute tanks with very few hardware failures. That said, my company chooses to have OEM support for replacing failed components and keeping firmware/bios/idrac updated. You can absolutely run these if you’re OK with 3rd party replacements or parting out spare machines. Some industries are more tolerant to this than others.
- sgarland 8mo agoI ran 3x R620s 24/7/365 in my homelab for ~6 years (well, other than when I moved, or shut one down for a clean-and-inspect, or lost power in excess of what my UPS could handle... thanks, Texas). The only things that failed during that time were a couple of sticks of RAM, and a PSU.
- sdbrown 8mo agoOn a related-different axis, I've consistently seen on-prem GPUs running identical workloads ~35% faster than the same workloads on the same cloud hardware, regardless of intermediate infra stack layering/versioning choices. Weird but I'm not complaining!
- jononor 8mo agoThey might be undervolting and underclocking the GPUs to improve longevity?
- scalemaxx 8mo agoEverything comes circle. Back in my day, we just called it a "data center". Or on-premise. You know, before the cloud even existed. A 1990s VP of IT would look at this post and say, what's new? Better computing for sure. Better virtualization and administration software, definitely. Cooling and power and racks? More of the same. The argument made 2 decades ago was that you shouldn't own the infrastructure (capital expense) and instead just account for the cost as operational expense (opex). The rationale was you exchange ownership for rent. Make your headache someone else's headache. The ping pong between centralized vs decentralized, owned vs rented, will just keep going. It's never an either or, but when companies make it all-or-nothing then you have to really examine the specifics.
- the_af 8mo agoAgreed. Also, a realistic assessment should not downplay the very real overhead and headache of managing your on-premise data center. It comes at a cost in engineering/firefighting hours, it's not painless. There's a reason this eternal ping pong keeps going on!
- adolph 8mo agoYeah, I think the major improvement of cloud services was the rationalization of them into services with a cost instead of "ask that person for a whatsit" and "hopefully the associate goomba will approve." All teams will henceforth expose their data and functionality through service interfaces https://gist.github.com/chitchcock/1281611 https://gist.github.com/chitchcock/1281611
- re-thc 8mo ago> you shouldn't own the infrastructure (capital expense) and instead just account for the cost as operational expense (opex) That was part of the reason. The real reason was the internal infrastructure team in many orgs got nowhere. There was a huge queue and many teams instead had to find infinite workarounds including standing up their own. The "cloud" provided a standardized way to at least deal with this mess e.g. single source of billing. > A 1990s VP of IT would look at this post and say, what's new? Speed. The US lives in luxury but outside of that it often takes a LONG time to get proper servers. You don't just go online. There are many places where you have to talk to a vendor with no list price and the drama continues. Being out of capacity can mean weeks to months before you get anywhere.
- siliconc0w 8mo agoYou can also buy the hardware and hire an IT vendor to rack and help manage it as smart hands so you never need to visit the datacenter. With modern beefy hardware, even large web services only need a few racks so most orgs don't even to manage a large footprint. Sure you have to schedule your own hardware repairs or updates but it also means you don't need to wrangle with the ridiculous cost-engineering, reserved instances, cloud product support issues or API deprecations, proprietary configuration languages, etc. Bare metal is better for a lot of non-cost reasons too, as the article notes it's just easier/better to reason about the lower level primitives and you get more reliable and repeatable performance.
- j45 8mo agoThat’s called managed servers or managed services. I have run bare metal and manage services you just have to be clear on what you have coverage for when disaster strikes or be willing to proactively replace hard drives before they die.
- lawrenceyan 8mo agoHetzner bare metal ran much of crypto for many years before they cracked down on it.
- regular_trash 8mo agoThe distinction between rent/own is kind of a false dichotomy. You never truly own your platform - you just "rent" it in a more distributed way that shields you from a single stress point. The tradeoff is that you have to manage more resources to take care of it, but you have much greater flexibility. I have a feeling AI is going to be similar in the future. Sure, you can "rent" access to LLM's and have agents doing all your code. And in the future, it'll likely be as good as most engineers today. But the tradeoff is that you are effectively renting your labor from a single source instead of having a distributed workforce. I don't know what the long-term ramifications are here, if any, but I thought it was an interesting parallel.
- rob_c 8mo agoAnd finally we reach the point where you're not shot for explaining if you invest in ownership after everything is over you have something left that has intrinsic value regardless of what you were doing with it. Otherwise, well just like that gym membership, you get out what you put into it...
- ex-aws-dude 8mo agoI can see how this would work fine if the primary purpose is for training rather than serving large volumes of customer traffic in multiple regions It would probably even make sense for some companies to still use cloud for their API but do the training on prem as that may be the expensive part.
- throwaway-aws9 8mo agoThe cloud is a psyop, a scam. Except at the tiniest free-tier / near free-tier use cases, or true scale to zero setups. I've helped a startup with 2.5M revenue reduce their cloud spend from close to 2M/yr to below 1M/yr. They could have reached 250k/yr renting bare-metal servers. Probably 100k/yr in colos by spending 250k once on hardware. They had the staff to do it but the CEO was too scared. Cloud evangelism (is it advocacy now?) messed up the minds of swaths of software engineers. Suddenly costs didn't matter and scaling was the answer to poor designs. Sizing your resource requirements became a lost art, and getting into reaction mode became law. Welcome to "move fast and get out of business", all enabled by cloud architecture blogs that recommend tight integration with vendor lock-in mechanisms. Use the cloud to move fast, but stick to cloud-agnostic tooling so that it doesn't suck you in forever. I've seen how much cloud vendors are willing to spend to get business. That's when you realize just how massive their margins are.
- re-thc 8mo ago> The cloud is a psyop, a scam. You're just young. > Suddenly costs didn't matter and scaling was the answer to poor designs. It did. Did you know that cloud cost less than what the internal IT team at a company would charge you? Let's say you worked on product A for a company and needed additional VM. Besides paperwork, the cost to you (for your cost center) would be more than using the company credit card for the cloud. > Sizing your resource requirements became a lost art In what way? We used to size for 2-4x since getting additional resources (for the in-house team) would be weeks to months. Same old - just cloud edition.
- throwaway-aws9 8mo ago> You're just young. And I feel great! > Did you know that cloud cost less than what the internal IT team at a company would charge you? Yes. Internal IT teams ran old-school are inefficient. And that's what the vendor tells you while they create shadow IT inside your company. Skip ITSM and ITIL... do it the SRE way. Until the cloud economist (real role) comes in and finds a way to extract more rent out of their customer base (like GCP's upcoming doubling rates on CDN Interconnect). And until internal IT kills shadow IT and regains management of cloud deployments. Cybersecurity and stuff... Back to square one. ITIL with cloud deployments. Some use cases will be way cheaper... but for your 100s of PBs of enterprise data, that's another story. And data gravity will kill many initiatives just based on bit movement costs. > Besides paperwork, the cost to you (for your cost center) would be more than using the company credit card for the cloud. To some extent. One is hard dollars the other is funny money. But I thought paying for cloud with the company credit card was a 2016 thing. Now it's paid through your internal IT cost center, with internal IT markup. I've seen petabytes of data move to the cloud and then we couldn't perform some queries on it anymore as that store wouldn't support it, and we'd need to spend 7 figures to move to another cloud database to query it. And that's hard dollars. Yes, during early cloud days it was lean and aimed at startups. Now it's aimed at enterprise, and for some reason lots of startups still think it's optimized for them. It's not and it hasn't been for a long time.
- epistasis 8mo agoAh Slurm, so good to see it still being used. As soon as I touched it in ~2010 I realized this was finally the solid queue management system we needed. Things like Sun Grid Engine or PBS were always such awful and burdensome PoS. IIRC, Slurm came out of LLNL, and it finally made both usage and management of a cluster of nodes really easy and fun. Compare Slurm to something like AWS Batch or Google Batch and just laugh at what the cloud has created...
- eubluue 8mo agoOn top of that, now when the US cloud act is again a weapon against EU, most European companies know better and are migrating in droves to colo, on-prem and EU clouds. Bye bye US hyperscalers!
- vadepaysa 8mo agoI was an on-prem maxi (if thats a thing) for a long time. I've run clusters that costed more than $5M, but these days I am a changed man. I start with PaaS like Vercel and work my way down to on-prem depending on how important and cost conscious that workload is. Pains I faced running BIG clusters on-prem. 1. Supply chain Management -- everything from power supplies all the way to GPUs and storage has to be procured, shipped, disassembled and installed. You need labor pool and dedicated management. 2. Inventory Management -- You also need to manage inventory on hand for parts that WILL fail. You can expect 20% of your cluster to have some degree of issues on an ongoing basis 3. Networking and security -- You are on your own defending your network or have to pay a ton of money to vendors to come in and help you. Even with the simplest of storage clusters, we've had to deal with pretty sophisticated attacks. When I ran massive clusters, I had a large team dealing with these. Obviously, with PaaS, you dont need anyone.
- majormajor 8mo agoIn addition to those sorts of non-first-hardware-purchase costs, the person writing the check needs to think long and hard about how bad an outage would be, and how much money it makes sense to budget simply to "avoiding outages." And the more important it is not to have any downtime, the more it's gonna cost to build up some sort of substitute for cross-datacenter cloud functionality. (You are also likely not going to be as good at either managing and configuring those networks, or hiring people to do so, as AWS, either.)
- cheema33 8mo ago> I was an on-prem maxi (if thats a thing) for a long time. I've run clusters that costed more than $5M, but these days I am a changed man. I have had a similar transformation. I still host non-critical services on-prem. They are exceptionally cheap to run. Everything else, I host it on Hetzner.
- b8 8mo agoSSD's don't last longer than HDDs. Also they're much more expensive due to AI now. They should move to cutdown on power costs.
- deleted 8mo ago[deleted]
- tgtweak 8mo ago>San Diego has a mild climate and we opted for pure outside air cooling. This gives us less control of the temperature and humidity, but uses only a couple dozen kW. We have dual 48” intake fans and dual 48” exhaust fans to keep the air cool. To ensure low humidity (<45%) we use recirculating fans to mix hot exhaust air with the intake air. One server is connected to several sensors and runs a PID loop to control the fans to optimize the temperature and humidity. Oh man, this is bad advice. Airborn humidity and contaminants will KILL your servers on a very short horizon in most places - even San Diego. I highly suggest enthalpy wheel coolers (kyotocooling is one vendor - switch datacenters runs very similar units on their massive datacenters in the Nevada desert) as they remove the heat from the indoor air using outdoor air (+can boost slightly with an integrated refrigeration unit to hit target intake temps) without switching the air from one side to the other. This has huge benefits for air control quality and outdoor air tolerance and a single 500KW heat rejection unit uses only 25KW of input power (when it needs to boost the AC unit's output). You can combine this with evaporative cooling on the exterior intakes to lower the temps even further at the expense of some water consumption (typically far cheaper than the extra electricity to boost the cooling through an hvac cycle). Not knocking the achievement just speaking from experience that taking outdoor air (even filtered + mixed) into a datacenter is a recipe for hardware failure and the mean time to failure for that is highly dependant on your outdoor air conditions. I've run 3MW facilities with passive air cooling and taking outdoor air directly into servers requires a LOT more conditioning and consideration than is outlined in this article.
- Torq_boi 8mo agoYes, it's easy to destroy the servers with a lot of dust and/or high humidity. But with filtering and ensuring humidity never exceeds 45% we've had pretty good results.
- kccqzy 8mo agoI remember visiting a small data center (about half the size of the Comma one) where shoe covers were required. Apparently they were worried about people’s shoes bringing in dust and other contamination.
- 8mo ago
- deadbabe 8mo agoClouds suck. But so does “on premises”. Or co-location. In the future, what you will need to remain competitive is computing at the edge. Only one company is truly poised to deliver on that at massive scale.
- deleted 8mo ago[deleted]
- 3acctforcom 8mo agoThe lowest grade I got in my business degree was in the "IT management" course. That's because the ONLY acceptable answer to any business IT problem is to move everything to the cloud. Renting is ALWAYS better than owning because you transfer cost and risk to a 3rd party. That's pretty much the dogma of the 2010s. It doesn't matter that my org runs a line-of-business datacentre that is a fraction of the cost of public cloud. It doesn't matter that my "big" ERP and admin servers take up half a rack in that datacentre. MBA dogma says that I need to fire every graybeard sysadmin, raze our datacentre facility to the ground, and move to AWS. Fun fact, salaries and hardware purchases typically track inflation, because switching cost for hardware is nil and hiring isn't that expensive. Whereas software is usually 5-10% increases every year because they know that vendor lock-in and switching costs for software are expensive.
- MagicMoonlight 8mo agoRight, but is that a like for like comparison? AWS has redundant data centres across the world and within each region. A file in S3 will never be lost, even if you store it for a thousand years. What happens if your city has a tornado and your data centre gets hit? Is your company now dead? And how much do you spend on all these sysadmins? 200k each? If you’re saving 20k/month by paying 100k/month in salaries, you aren’t saving anything.
- 0xbadcafebee 8mo agoIf your business relies on compute, and you run that compute in the cloud, you are putting a lot of trust in your cloud provider. Cloud companies generally make onboarding very easy, and offboarding very difficult. If you are not vigilant you will sleepwalk into a situation of high cloud costs and no way out. If you want to control your own destiny, you must run your own compute. This is not a valid reason for running your own datacenter, or running your own server. Self-reliance is great, but there are other benefits to running your own compute. It inspires good engineering. Maintaining a data center is much more about solving real-world challenges. The cloud requires expertise in company-specific APIs and billing systems. A data center requires knowledge of Watts, bits, and FLOPs. I know which one I rather think about. This is not a valid reason for running your own datacenter, or running your own server. Avoiding the cloud for ML also creates better incentives for engineers. Engineers generally want to improve things. In ML many problems go away by just using more compute. In the cloud that means improvements are just a budget increase away. This locks you into inefficient and expensive solutions. Instead, when all you have available is your current compute, the quickest improvements are usually speeding up your code, or fixing fundamental issues. This is not a valid reason for owning a datacenter, or running your own server. Finally there’s cost, owning a data center can be far cheaper than renting in the cloud. Especially if your compute or storage needs are fairly consistent, which tends to be true if you are in the business of training or running models. In comma’s case I estimate we’ve spent ~5M on our data center, and we would have spent 25M+ had we done the same things in the cloud. This is one of only two valid reasons for owning a datacenter, and one of several valid reasons for running your own server. The only two valid reasons to build/operate a datacenter: 1) what you're doing is so costly that building your own factory is the only profitable way for your business to produce its widgets, 2) you can't find a datacenter with the location or capacity you need and there is no other way to serve your business needs. There's many valid reasons to run your own servers (colo), although most people will not run into them in a business setting.
- ynac 8mo agoNot nearly on the article's level, but I've been operating what I call a fog machine (itsy bitsy personal cloud) for about 15 years. It's just a bunch of local and off-site NAS boxes. It has kinda worked out great. Mostly Synology, but probably won't be when their scheduled retirement comes up. The networking is dead simple, the power use is distributed, and the size of it all is still a monster for me - back in the day, I had to use it for a very large audio project to keep backups of something like 750,000 albums and other audio recordings along with their metadata and assets.
- barbazoo 8mo agoAnd now go do that in another region. Bam, savings gone. /s What I mean is that I'm assuming the math here works because the primary purpose of the hardware is training models. You don't need 6 or 7 nines for that is what I'm imagining. But when you have customers across geography that use your app hosted on those servers pretty much 24/7 then you can't afford much downtime.
- farceSpherule 8mo ago[dead]
- pelasaco 8mo agoif i understood correctly, you dont kubernetes, rights? Did you consider it?
- stego-tech 8mo agoIT dinosaur here, who has run and engineered the entire spectrum over the course of my career. Everything is a trade-off. Every tool has its purpose. There is no "right way" to build your infrastructure, only a right way for you. In my subjective experience, the trade-offs are generally along these lines: * Platform as a Service (Vercel, AWS Lambda, Azure Functions, basically anything where you give it your code and it "just works"): great for startups, orgs with minimal talent, and those with deep pockets for inevitable overruns. Maximum convenience means maximum cost. Excellent for weird customer one-offs you can bill for (and slap a 50% margin on top). Trade-off is that everything is abstracted away, making troubleshooting underlying infrastructure issues nigh impossible; also that people forget these things exist until the customer has long since stopped paying for them or a nasty bill arrives. * Infrastructure as a Service (AWS, GCP, Azure, Vultr, etc; commonly called the "Public Cloud"): great for orgs with modest technical talent but limited budgets or infrastructure that's highly variable (scales up and down frequently). Also excellent for everything customer-facing, like load balancers, frontends, websites, you name it. If you can invoice someone else for it, putting it in here makes a lot of sense. Trade-off is that this isn't yours, it'll never be yours, you'll be renting it forever from someone else who charges you a pretty penny and can cut you off or raise prices anytime they like. * Managed Service/Hosting Providers (e.g., ye olde Rackspace): you don't own the hardware, but you're also not paying the premium for infrastructure orchestrators. As close to bare metal as you can get without paying for actual servers. Excellent for short-term "testing" of PoCs before committing CapEx, or for modest infrastructure needs that aren't likely to change substantially enough to warrant a shift either on-prem or off to the cloud. You'll need more talent though, and you're ultimately still renting the illusion of sovereignty from someone else in perpetuity. * Bare Metal, be it colocation or on-premises: you own it, you decide what to do with it, and nobody can stop you. The flip side is you have to bootstrap everything yourself, which can be a PITA depending on what you actually want - or what your stakeholders demand you offer. Running VMs? Easy-peasy. Bare metal K8s clusters? I mean, it can be done, but I'd personally rather chew glass than go without a managed control plane somewhere. CapEx is insane right now (thanks, AI!), but TCO is still measured in two to three years before you're saving more than you'd have spent on comparable infrastructure elsewhere, even with savings plans. Talent needs are highly variable - a generalist or two can get you 80% to basic AWS functionality with something like Nutanix or VCF (even with fancy stuff like DBaaS), but anything cutting edge is going to need more headcount than a comparable IaaS build. God help you if you opt for a Microsoft stack, as any on-prem savings are likely to evaporate at your next True-Up. In my experience, companies have bought into the public cloud/IaaS because they thought it'd save them money versus the talent needed for on-prem; to be fair, back when every enterprise absolutely needed a network team and a DB team and a systems team and a datacenter team, this was technically correct. Nowadays, most organizational needs can be handled with a modest team of generalists or a highly competent generalist and one or two specialists for specific needs (e.g., a K8s engineer and a network engineer); modern software and operating systems make managing even huge orgs a comparable breeze, especially if you're running containers or appliances instead of bespoke VMs. As more orgs like Comma or Basecamp look critically at their infrastructure needs versus their spend, or they seriously reflect on the limited sovereignty they have by outsourcing everything to US Tech companies, I expect workloads and infrastructure to become substantially more diversified than the current AWS/GCP/Azure trifecta.
- MagicMoonlight 8mo agoFor ML it makes sense, because you’re using so much compute that renting it is just burning money. For most businesses, it’s a false economy. Hardware is cheap, but having proper redundancy and multiple sites isn’t. Having a 24/7 team available to respond to issues isn’t. What happens if their data centre loses power? What if it burns down?
- Hasz 8mo agoThis is hackernews, do the math for the love of god. There are good business and technical reasons to choose a public cloud. There are good business and technical reasons to choose a private cloud. There are good business and technical reasons to do something in-between or hybrid. The endless "public cloud is a ripoff" or "private clouds are impossible" is just a circular discussion past each other. Saying to only use one or another is textbook cargo-culting.
- wessorh 8mo agowhat is the underling filesystem for your kv store, it doesn't appear to use raw devices.
- dh2022 8mo agoLOL’ed IRL at “ In a future blog post I hope I can tell you about how we produce our own power and you should too.” Producing own power as a pre-requisite for running on-prem is a non-starter for many.
- asdfman123 8mo ago"It's really not hard to create your own coal power. Our engineers have built a small coal power generator and simply get coal from our mines (which I'll describe in a future blog post)."
- dh2022 8mo agoLOL'ed again IRL :).
- sakopov 8mo agoDoes anyone remember how cloud prices used to trend down? That was about 6 years ago and then seemingly after the pandemic everything started going the other way.
- swordsith 8mo agoRecently learned about tailscale and have been accessing my project from my phone, It's been a game changer. The fact that they support teams of up to 3 people and 100 devices on the free plan is awesome imo. Running locally just makes me feel so much more comfortable.
- alecco 8mo agoCounterpoint: "Why I'm Selling All My GPUs" https://www.youtube.com/watch?v=C6mu2QRVNSE https://www.youtube.com/watch?v=C6mu2QRVNSE TL;DW: GPU rental arbitrage is dead. Regulation hell. GPU prices. Rental price erosion. Building costs rising. Complexity of things like backup power. Delays of connection to energy grid. Staffing costs.
- SomaticPirate 8mo agoThis makes sense for HPC and ML workloads. Big batch jobs where you are pushing the hardware and having everything local is a clear advantage. Also this company sells hardware so it makes sense for them to have hardware experience. Still think that for the majority on here, needing to make a physical phone call to their data center team (!!) to rack a server is a nutty proposition. You think the AWS api is slow? Trying calling Steve. If you have fixed compute costs after a year, sure, look at pulling some stuff on prem.
- BLACKCRAB 8mo ago[dead]
- yawnxyz 8mo agorunning your own ai inference is quite stressful, and reading this article definitely makes me feel stressed