15 ms·
Ahrefs saved $400m in 3 years by not going to the cloud
- pauldambra 4y agoI really don't miss running my own kit (colocated or directly owned) I don't miss cycling around Manchester on a Bank Holiday weekend because I'd miscalculated how much network cabling I'd need for an upgrade. I don't miss keeping a spreadsheet of storage so I knew when to order disks, negotiating with suppliers for cost of new disks because I was buying a slightly smaller bulk than AWS I don't miss having to explain to folk in datacenter support that they could take the disks out of my failed server and put them in a new server if they had one available I don't miss the day the single point of failure in the rack failed and everything was offline while I waited for a new doohicky to be shipped to me because it didn't make sense to keep spares of everything on hand I don't miss trying to figure out if some new generation of server hardware would work for or would fit in my rack as manufacturers stopped making the kit we did use I'm not going to say every workload should run in the cloud (cliche nod to StackOverflow) but it certainly isn't free to get all of the benefits
- cs702 4y agoI recently heard from a senior IT person that the company for which he works could save many millions by moving a certain application from a hairball of micro-services at AWS to a much simpler architecture on colocated servers... but the company's management doesn't want to hear any of it. In fact, management wants IT to move every legacy application that's not yet on the cloud to the cloud, specifically to AWS, because it's considered "more flexible" and perceived as "the future." He's taken to calling AWS the "lifetime employment program for enterprise software developers."
- mirochnik 4y agoThis is a new type of conflict of interests. I assumed only good intentions from IT guys. But you're right that some of them may be more interested in complexity of used solutions, bigger budgets, and their personal job security than in company results and sustainability.
- fidgewidge 4y agoIt's not like they want complexity, it's the latter: the "nobody ever got fired for buying AWS" problem. They stated it explicitly. They're terrified of anything counter-cultural or that seems to be going against the tide because they want to be seen as in the future and not stuck in the past. Fashion, FOMO and weakness combine to yield a "cloud at any cost" mentality.
- mirochnik 4y agoWow! It sounds even worse.
- bombcar 4y agoA significant number of "cloud migrations" involve moving an application running on Linux on a server to running on Linux on a VPS for similar or higher cost. Not even using ANY of the "features" of the cloud, just treating it like a server. But then they can say "cloud" and everyone's happy.
- hotpotamus 4y agoAh yes, the "actual" way to use cloud is to undergo a major re-write of your application such that it is inextricably tied into the proprietary PaaS options of your cloud provider. Cloud is amazing for highly variable loads. I also see it as a bit of a luxury service for ops types (like me) - I don't have to go to a DC and manage hardware or deal with a DCops crew, so that's nice, but you probably wouldn't buy luxury cars if you were starting a courier business.
- bombcar 4y agoWhat really annoys me is that few if any clouds actually allow "dynamic scaling" of a single instance running operating system where you can hot-add more RAM or more CPU, with or without a restart. Some can do it, some require you to basically image the entire machine onto a new one, some can't handle it at all.
- amluto 4y agoI’ve never worked at a cloud company, but I know more than I’d really like about some of the stuff under the hood. First, changing out the CPU type to something that isn’t almost identical out from under a running VM is a mess and may require degrading the system by removing features from the starting CPU. Switching manufacturers at runtime (AMD vs Intel), while sort of possible in theory, is effectively a lost cause. To add CPUs, you have to deal with the architecture’s nasty hot-add CPU mechanism, and the guest OS needs to be expecting those CPUs starting from when it boots, and that expectation isn’t free. (The latter is Linux’s num_possible_cpus vs num_present_cpus.) Adding memory involves getting that memory into the kernel’s memory map in the right places. This is more complex than one would like. Removing memory is worse. And all this happens, in the cloud data center, on essentially normal hardware. If you want to add RAM, the CPUs need to be on a system with more RAM, preferably on the same NUMA node. And vice versa for adding CPUs. If the tenant is paying for local storage, that needs to move, too.
- gadders 4y agoI think a lot of companies are going to find that they have the same relationship with AWS as banks do with Bloomberg i.e. a massive spend that they constantly try to manage downwards but can never escape.
- bfeynman 4y agoFrom a biz perspective, that can be a bad move and signals that you're not in growth phase but in cutting costs to preserve profits. Using cloud derisks you for future costs because it's managed for you, and then you allocate capex towards making better features. Doing it locally means having a team of IT admins/professionals you normally only need a few of which also incurs a lot of cost and risk.
- hardware2win 4y agoLets replace 4 cheap admins with 3 expensive devops engs Yaay savings!
- amluto 4y agoSiloing expenses into capex and opex gets a bit silly when the cost effectiveness of that opex becomes low enough. For some cloud services, one can replace it with on-prem or colo’d equipment and recover the capex in months. Similarly, “easy” scalability doesn’t necessarily make up for COGS when scaling up. If I can flip the autoscale switch while I grow to $1bn ARR and pay $750M/yr to the cloud, I’m not doing nearly as well as if I hire ten FTEs and pay $150M in combined capex and opex over two years. (And if my COGS is 75% of the list price, it’s hard to give discounts, pay for sales, etc)
- freitzkriesler2 4y agoTale as old as time. Heard the same with exchange, MSP contracts, etc etc some suit somewhere always has some CYA reason no matter how silly or wrong it may me.
- warent 4y agoWhen consulting for a seed funded startup, I suggested they buy their own servers and colocate them. It could save them almost 100k a year. Everyone looked at me like an alien speaking a different language, lol. Then was politely dismissed, even though I'm experienced in running hardware. AWS, GCP, Azure, really managed to expertly pull off the greatest heist of all time. They're useful for sure, but somehow they've gotten everyone to drink this potent marketing kool-aid that if you're not using them then you're not serious about success. So cringe.
- senttoschool 4y agoThey'd have to keep you employed to manage the hardware. What happens if you get run over by a bus (which I never hope will happen)? What happens if you decide to move to another city? There are a lot more "what ifs" by employing you to build and manage the hardware. Saving $100k would not be worth the hassle/risk for most well funded startups.
- warent 4y agoAgain this is the kool-aid speaking. The co-location manages most of the complexity and the probability of failure is basically zero for the first 4 years. Hardware these days is really good. People have this illusion that servers are extremely hard to maintain because big companies constantly have to maintain their 1000s, and because clouds are incentivized to sell this lie. It would still be a huge cost saving and trivial for them to contract someone to maintain the servers for like 1 hour a month with emergency on-call. Of which there are numerous such people who are not me.
- senttoschool 4y agoI used to build PCs. You can get the same hardware for nearly half the cost of pre-built PCs like HP or Dell. I did a lot of research on the best bang for the buck parts and the most reliable parts. If something broke, just pull it out and replace it. Easy. I worked in an IT department as an intern. I thought this company should build custom PCs for all their workers. You can get a much faster CPU or much faster GPU for the same price! It's a no brainer! Then I realized that if I wanted to build custom PCs for all the workers, I'd have to provide support for all of them. I'd have to be as good as CDW, who was supplying the company with PCs. This means next-day repairs and replacements. Drivers. Security. Firmware updates. For hundreds of PCs. Needless to say, a lot of people under appreciate all the things that vendors like AWS or CDW provide. They think they can do it because they know how it's done. But they're ignoring a lot of other factors that can come back to bite you. I'm sure you're very competent at managing server hardware. But if I'm a seed stage startup, I'd politely decline your offer.
- ejb999 4y agoBut that doesn't really prove much does it? You are saying they could rearchitect/simplify the whole thing and move it on site to save money - but what about rearchitecting the whole thing and leaving it on AWS? Which part would actually saving them the money? - rearchitecting/simplifying, or changing where it runs? There would need to be an apples-to-apples comparison to be meaningful. That said, I don't believe everything needs to be in the cloud - but maintaining data-centers are also very expensive, and need to be accounted for very differently.
- sshine 4y agoI helped migrate a PHP app to AWS and K8s for a startup that got acquired a couple of years ago. Before it was running on a dedicated Xeon with the database available on a UNIX socket. After... where do I start. The app was very database heavy and experienced timeouts due to too many database calls (IOPS). Not a problem prior to migration, because you don’t pay separately for UNIX sockets. Just the IOPS ended up costing several hundred bucks a month. All environment variables were stored as secrets at $.40/mo. per variable. I know this is peanuts, but it is so illustrative of the pricing model. It feels like those candy bags where each tiny piece of candy is individually wrapped in plastic. I found $35k/mo I’d have spent differently.
- oxfordmale 4y agoIf you store secrets in the Parameter Store, there is no cost per month. In AWS you need to be careful what you use.
- oxfordmale 4y agoOn-Premise Indirect (Hidden) Costs There are a variety of other expenses associated with an on-premise environment considered indirect expenses. These expenses are often referred to as “hidden” expenses due to how often they are overlooked rather than “hidden.” These include: The real estate of the storage space used for the servers Tools used for temperature control in the data center The cost of set up, configuration, and ongoing upgrades Staff salaries for administrators that maintain an on-premise data center Networking infrastructure set up and ongoing maintenance The cost of downtime while the team troubleshoots the issue Productivity lost when the system experiences downtime The cost of keeping the servers powered 24/7 Depreciation of the hardware and software Time spent on disaster recovery Administrative costs associated beyond IT staff such as HR, purchasing, financing, and other departments. https://pmsquare.com/analytics-blog/2022/6/9/calculating-on-prem-vs-cloud-costs https://pmsquare.com/analytics-blog/2022/6/9/calculating-on-...
- fidgewidge 4y agoYou speak like clouds never have downtime, or don't charge you for space and power. Look at the article - that one company would needed to have paid more than their entire revenue to AWS for worse hardware. They'd have spent $400M in 3 years on the cloud! There are no "hidden costs" that can even begin to approach a fraction of that. As for staff salaries, lol. It's not like AWS is self administering.
- oxfordmale 4y agoIn the cloud, costs are transparent, as you receive a single monthly bill. When you self-host, there are many hidden expenses you would need to track down (energy, admin, maintenance, etc.) The figures in the article are for illustration purposes only and should be taken with a large pinch of salt. The author doesn't detail what hardware they are running, what EC2 instances he has selected for comparison and how comparable the storage statistics are. I would also love to hear the Finance Department's version of his calculations. AWS is more expensive than self-hosting. However, it is not as skewed as the author claims. Otherwise, very few companies would be using the cloud.
- j45 4y agoRetail cloud setups provide standardization that hosting on bare metal may not. The irony of it is most of the cloud services are open-source technologies packaged up to be easy to use and administer from a web interface. Maybe if someone came up with a set of tooling in AWS to use many lightsail VMs to provide more of the basic AWS services, it could be a way to get the best of both worlds.
- Hnrobert42 4y agoWhat if they went from a hairball in AWS to a much simpler architecture in AWS. Wouldn’t that be a better comparison?
- VWWHFSfQ 4y agoI often fantasize about moving a bunch of my company's crap off of ECS/EKS and onto a managed colo with just old-school ansible deployments. I've even spec'd out some bare metal from our local colocation facility and the servers they offer are so ridiculously powerful and CHEAP. We could run our entire production system in a half-rack for about 10% of what we pay AWS. But alas, nobody will ever go for it and I'm left to dream.
- tiffanyh 4y ago> "We could run our entire production system in a half-rack for about 10% of what we pay AWS" Performance. You'd also get way faster hardware too.
- glintik 4y agoAnd very wide choice of hardware and tuning abilities. Also network config can speed up overall performance greatly.
- Nextgrid 4y agoDirect-attach storage is a major improvement for things like databases, etc which you pretty much never have in the cloud, at least not in a persistent form.
- glintik 4y agoDon’t agree. Multiple database nodes with enterprise grade SSDs is a way better in many terms.
- Nextgrid 4y agoFor conventional relational databases like Postgres, multiple nodes only give you reliability, not performance (ignoring things like read-only replicas which your application explicitly has to choose to query), so horizontal scaling doesn't help there. Enterprise-grade SSDs is what I meant by direct-attach storage - I was comparing against network storage which is what all cloud providers use (your EBS volume is accessed over the network internally, which incurs some latency, ultimately limits IOPS and destroys random access performance which can't be cached or read-ahead by the underlying hypervisor).
- deleted 4y ago[deleted]
- blakesterz 4y agoPretty interesting point there at the end of the post: But with the mass layoffs in Big Tech in recent months, this may be an opportunity to re-evaluate the approach to the cloud, consider a reverse migration from the cloud, and hire seasoned professionals of the data center world. I do think with all these layoffs we can't help but see some interesting things happen from people who have been let go. Either through their work at other places, or new startups.
- bfeynman 4y agodespite what everyone thinks that doing things locally is cheaper, if you're running a business and allocating money towards headcount and benefits plus additional risk of maintaining team it rarely makes sense. HN is full of people who run a hobby server and think that a fortune 500 should manage its own PostGres
- steve1977 4y agoYou still need headcount to manage systems on AWS. And not really fewer then if you would run them locally. Source: managed systems locally, am now managing systems on AWS
- mlnj 4y agoSurely the cloud managed solution reduces the amount of custom magic that I have to put into place to keep things running. Everything is standardized & documented. There is a huge community to turn to if help is needed. And very less work is required if I get hit by a bus and someone else is now the owner of the huge system that I previously owned.
- endisneigh 4y agobut you can run kubernetes on bare metal
- steve1977 4y agoIt is standardized and documented if you DO document it and properly use IaC like CloudFormation or Terraform. But you can absolutely have people do ad-hoc changes in AWS that are not gonna be reproducible. Also, all of your points apply to non-cloud systems too.
- antondd 4y agoEntirely predictable. Cloud is good for some use cases. Self hosting is good for some other use cases. Neither is objectively “better”. Both should be seen and used in the context of a task in hand. But alas it is human nature to pick up a hammer, call it the best thing in the world, and start whacking it at every problem. See: cloud, blockchain, k8s, and now chatgpt/ai.
- antimius3000 4y agoIn the last 15 years that I've seen this type of clickbait, 100% of the time it's written by someone who doesn't know how to make a TCO calculation. Of course, this time was no different. Should have saved myself the agony after reading the first sentence: "Clouds for IT infrastructure are so popular lately that moving into the cloud has become a trend.". Please.
- nik736 4y agoSo let us know what is wrong and how a proper TCO calculation looks like. What's missing?
- dtech 4y agoAnything that you have to do "extra" compared to managing the hardware yourself. E.g. this article is missing even basic stuff like the (prorated) salary costs of employees buying, installing and servicing the hardware.
- humanistbot 4y ago> E.g. this article is missing even basic stuff like the (prorated) salary costs of employees buying, installing and servicing the hardware. You say that like you don't need a dedicated team to manage AWS.
- OtomotO 4y agoAh, so now the developers will service the AWS instances from your end which means they can work less on delivering new features... also left out of the equation quite often :)
- sokoloff 4y agoThey’ll also wait less for IT to service their tickets requesting new infrastructure, which leads to more new feature development. There are a variety of trade-offs; not all of them have the same sign.
- 4y ago
- deleted 4y ago[deleted]
- bfeynman 4y agoThis napkin math is pointless, if you're a business leader, you're going to factor in non physical costs as well, since you need a team to run hundreds of servers most likely. If business all of a sudden starts to go under, you can pull plug on cloud, but you'd have to write it off if self hosting. Hosting yourself makes sense if you're providing hardware level services like storage or compute. In that case, going to the cloud is literally financing your (potential ) competitor.
- steve1977 4y agoYou still need a team to manage the stuff on AWS. Probably the team needs to be even bigger. If you’re a business leader, you‘re probably just blindly following what your peer business leaders are doing, that’s why you’re on AWS (in most cases at least)
- Nextgrid 4y agoProviders like OVH, Hetzner, etc will provide you servers and handle maintenance just like cloud providers, and it turns out that in both cases you need to hire sysadmins (or "DevOps" engineers). You don't need extra staff unless you do actual on-prem which very few companies need (most can be served by a provider like the aforementioned ones).
- botsbreeder 4y agoDid you see how much we saved comparing to AWS? We could even hire an admin for each of 850 servers and would still have money left. But we have only one person taking care of hardware replacements. That is enough for the whole setup.
- Dunedan 4y agoSorry, but if I'm reading the article correctly you didn't save any money, because you weren't using AWS in the first place. All you did was to do some back of the napkin math what it might cost you to run something on AWS.
- 4y ago
- bluedino 4y ago> We use high core-count CPUs, 2TB RAM, and 2x 100Gbps per server. On average, our servers have about 16x 15TB drives. > So let’s say we run our 850 servers I'm not familiar with this company, but it says on their website We’ve been crawling the web for over 10 years, collecting and processing petabytes of data every day., they have their own search engine (named 'Yep), provide all kinds of reports, etc. I'm guessing they are running huge Hadoop clusters or something? One of those workloads that just isn't suited for the cloud, and they're just taking the advice of Joel Spolsky “If it's a core business function — do it yourself, no matter what.”
- debarshri 4y agoI'm familiar with the product. They help you optimize your SEO, provider serp scores, backlink health, organic keywords performance etc. Basically, everything to do with Google SEO. I think calling them "just a bunch of hadoop cluster" might be disservice to the amount of effort that goes into building a product like this esp. When millions of digital marketing professionals use their service.
- KyeRussell 4y agoI feel like you’ve used a dishonest interpretation of this comment to push back against it and flex that you know what this company does? The only person ther said “just” is you. And “large Hadoop clusters” can be the most resource-intensive component of their infrastructure, but not necessarily the company’s value-add, or where most of their effort goes. Stop being adversarial.
- lynx23 4y agoI wonder how many wikileaks@paypal alike events it will need until companies understand that they make themselves far too exposed to total-takedown when moving everything to the cloud.
- oxfordmale 4y agoI feel this is comparing apple and pears. For sure AWS is more expensive, however, you looking at a markup of 50%, at most 100%.
- botsbreeder 4y agoWould be 1000% in our comparison
- oxfordmale 4y agoYour calculation is over simplified. For example, hardware just doesn't run for five years without glitch once once you have multiple servers. For sure Cloud is more expensive, as it charges a premium for not having to invest upfront in a lot of hardware, however,if it was 1000% no one would be using the Cloud.
- Symbiote 4y agoScaling how often I have failures on my 80 servers, of which 50 are similar in spec to Ahrefs (16 HDDs), it should still be less than one full-time position to manage 850 servers. Probably much less, as I'm a bit inefficient when it's <5% of my job. I have a fixed day or two each year working out what to buy and getting quotes (which would be the same time whether it's 20 or 200 servers), half a day per 20 servers for installation in racks, and call it half a day per 20 for initial configuration (most of the time to script the process, so it would also be similar for 200). After that, it's at absolute most an hour every couple of months to deal with a failed disc or similar.
- altdataseller 4y agoOh, that's not even close. It's closer to 500%
- oxfordmale 4y agoBusinesses are crazy but not that crazy. There are a number of hidden costs when self hosting. For sure Cloud is more expensive, but if you quote 500%, you are overlooking hidden costs.
- bastawhiz 4y agoAhrefs could probably save another few hundred million if they didn't repeatedly visit the same links over and over indefinitely to find the same error codes or a binary file that they surely don't care about. I see them in my logs for my podcast hosting service, hitting the same 404s for months and months. They end up hitting audio files and downloading (or attempting to download) many gigabytes of content each day. They don't do anything with audio! As soon as they get the HTTP headers, they could say "oops, I don't care about this" and disconnect. And in many cases, they got those URLs from the enclosure tags of RSS feeds... I'm not sure what they were expecting! In the cloud, not in the cloud, frankly it seems to me like their business would be far more efficient if they tuned what their crawler actually crawled.
- lazyfanatic 4y agoNo, throw hardware at it! My life for the last 25 years has been trying to optimize code via I/O. My favorite design pattern is a for loop that tries once per item and completely fails to never try again, nor report it failed at all. /s
- drcongo 4y agoAhrefs is like my white whale. I hate them. I can't stop myself from battling them. I hate the fact that they ignore status codes, I hate the fact that they ignore crawl rate specs in robots.txt, I hate the fact that they crawl URLs on my sites that have never and will never exist and keep coming back to do it again.
- dclusin 4y agoThey haven’t included an amortized cost breakdown with depreciation of hardware including backup. Or the geographically local personnel required (data center remote hands or dedicated employee that gets the call?). Even still I suspect it’s still probably cheaper. Also it looks like these are retail prices. You can get cost savings by negotiating but I’m sure there’s an NDA required.
- manv1 4y agoIt's good they did the math. AWS doesn't fit everything. Really, the main things that cost actual money in AWS are memory and CPU (that's including elasticache, RDS, etc). Bandwidth can be negotiated away. I'm sure you could do some kind of super deal on CPU/RAM, but we've never bothered. Would it be worth it to rearchitect your app so it doesn't use 2TB of RAM? Probably not. You have a lot of sunk costs, and redoing everything will probably break everything. You guys are big enough that CapEx doesn't matter that much. You just need big boxes with lots of bandwidth. If you're happy with what you've got, stay with it.
- DaveExeter 4y agoI block Ahrefs via IPtables. hydrogen033-ext2.a.ahrefs.com hydrogen042-ext2.a.ahrefs.com hydrogen106-ext2.a.ahrefs.com hydrogen326-ext2.a.ahrefs.com
- bediger4000 4y agoI kind of wish they hadn't. Ahrefs claims to be in the "SEO" business, which as far as I'm concerned, has helped ruin web search, and indeed, the culture of the web in general. Also, Ahrefs bot doesn't handle some things very well. I made an "infinite web site" in PHP a while back, and used Apache mod_rewrite to send every Ahrefs request to the infinite web site PHP program: https://github.com/bediger4000/infinite-fake-website https://github.com/bediger4000/infinite-fake-website Ahrefs bot really freaked out, unlike some professional bots like Google's, and even Yandex' bot.
- KomoD 4y ago> Ahrefs bot really freaked out, unlike some professional bots like Google's, and even Yandex' bot. > Also, Ahrefs bot doesn't handle some things very well. But that was... 7 years ago?
- bediger4000 4y agoDo you really think they act any differently? They're in the SEO business, which is directly parasitic. They're not going to spend money improving their software, that's a cost center, not a money maker.
- sebiandev 4y agoIt depends. Are these PAYGO prices or reserved instance prices? There is a HUGE difference. If these are PAYGO prices, the cost would probably be about the same to have your infra reserved in AWS WITH the added flexibility. This is why not just ANYONE should be able to spin up infrastructure in cloud providers. You should hire someone who knows what they're even doing.
- henriquez 4y agoHe compared against AWS 3 year reserve pricing.
- manv1 4y agoOne interesting question is: if you started today, would you be able to afford boxes that cost 60k? Or would you do your software some other way that doesn't require 2TB of RAM? Obviously when you started your stuff didn't need 2TB of RAM. If I read your history correctly I don't think you could even buy a box with 2TB of RAM back then. That's enterprise grade hardware, which today costs a fortune. Back in the day it would be a bigger fortune. Instead, you probably started the way everyone else did, with maybe a 4GB or 8GB linux box at home then built things up from there and a bunch of curl scripts and a local instance. So why the resource requirement? An in-ram database? Mainly curious.
- Symbiote 4y agoIt says they have 850 servers, so 1700TB RAM (if they're all like this). If it's a big Hadoop cluster or similar, there's potentially huge performance gains by keeping data in-RAM during processing. They have many petabytes of data, so I doubt the type of processing they are doing today would have been possible when they first started.
- ThePowerOfFuet 4y agoHere's the tracking-free version of OP's link: https://tech.ahrefs.com/how-ahrefs-saved-us-400m-in-3-years-by-not-going-to-the-cloud-8939dd930af8 https://tech.ahrefs.com/how-ahrefs-saved-us-400m-in-3-years-...