20 ms·
AWS to bare metal two years later: Answering your questions about leaving AWS
- alyxya 1y agoWith AI making it possible to use natural language to modify code, bare metal can make things easier to use with your own code and customization. Abstractions tend to be harder to reason about and have more limited functionality in exchange for being easier to get started on some standard setup.
- JCM9 1y agoFor smaller operations I’d still go with a rent-a-server model with AWS. Theirs is a critical mass though where roll your own makes sense. The long term app model on the market model is shifting much more towards buying services vs renting infrastructure. It’s here where the AWS case falls apart with folks now buying Planet Scale vs RDS, buying DataBricks over the mess that AWS puts for for data lakes, working with model providers directly vs the headaches of Bedrock. The real long term threat is AWS continues to whiff on all the other stuff and gets reduced to a boring rent-a-server shop that market forces will drive to be very low margin. Yes a lot of those 3rd party services will run on AWS but the future looks like folks renting servers from AWS at 7% gross margin and selling their value-add service on top at 60% gross margin.
- cmiles8 1y agoA bunch written about this recently by analysts. That is the “bear” outlook on AWS
- ramon156 1y agoThis doesn't really explain why you wouldn't just get a hetzner. I don't have much experience with either, but if you know how to setup your infra then hetzner seems like a no-brainer? I do not want to be tied to AWS where I have no idea what my bill will be
- JCM9 1y agoDepending on the use case you very much could just use Herzner. A simpler and more transparent customer experience than trying to navigate the mass complexity of AWS for basic stuff.
- pbiggar 1y ago[flagged]
- pbiggar 1y agoI'll add that Amazon execs told us to our face that they use your AWS usage to decide if they were going to copy your services (which they did in our case), so any dev or ops tool on AWS shouldn't be there anyway for competitive reason.
- mstipetic 1y agoWhat was the service they cloned from you?
- pbiggar 1y agoCircleCI - their version is CodeBuild. Note: edited the comment above to use "copy" instead of "clone" to be more accurate.
- aitchnyu 1y agoRelated, do they check your usage of their proprietary tools (ease of leaving) to negotiate your bills?
- al_borland 1y agoThey do this in the Amazon store as well, making their own knock off versions of popular products. Very interesting that it extends into the digital space as well. This really feels like it should be illegal, and if not, I’m not sure how there hasn’t been a massive revolt against them by people who make and sell things. If they are too big and powerful to revolt against, that sounds like a monopoly that needs to be broken up.
- marcinzm 1y agoI suspect if you look at your full supply chain for your bare metal the people they also sell to will not make you happy either.
- nik736 1y agoIt's an interesting article, thanks for that. What people forget about the OVH or Hetzner comparison is that for those entry servers they are known for, think the Advance line with OVH or AX with Hetzner. Those boxes come with some drawbacks. The OVH Advance line for example comes without ECC memory, in a server, that might host databases. It's a disaster waiting to happen. There is no option to add ECC memory with the Advance line, so you have to use Scale or High Grade servers, which are far from "affordable". Hetzner per default comes with a single PSU, a single uplink. Yes, if nothing happens this is probably fine, but if you need a reliable private network or 10G this will cost extra.
- jammo 1y agoYes, but there are options for dedicated server providers who offer dual PSU and ECC ram etc. It's more expensive though for e.g a 24 Core Epyc with 384GB RAM dual 10G netowork is like $500/month (though there's smaller servers on serversearcher.com for other examples)
- vjerancrnjak 1y agoIs there software that works without ECC RAM ? I think most popular databases just assume memory never corrupts .
- torginus 1y agoI'm pretty sure they keep internal internal checksums at various points to make sure the data on disk is intact - so does the filesystem, I think they can catch when memory corruption occurs, and can roll back to a consistent state (you still get some data loss). But imo, systems like these (like the ones handling bank transaction), should have a degree of resiliency to this kind of failure, as any hw or sw problem can cause something similar.
- lossolo 1y agoThese concerns are exaggerated. I've been running on Hetzner, OVH and friends for 20 years. During that time I've had only two issues, one about 15 years ago when a PSU failed on one of the servers, and another a few years ago when an OVH data center caught fire and one of the servers went down. There have been no other hardware issues. YMMV.
- aetherspawn 1y agoOk but what about a dedicated OVH for example? Those are about 70% cheaper than AWS, so is it still worth it to colo?
- bilekas 1y agoDid you read the article ? The main point of this and the prior article is that YES colocation/baremetal IS a better option for this company (and I would argue the majority of AWS users) reference : https://news.ycombinator.com/item?id=38294569 https://news.ycombinator.com/item?id=38294569
- bilekas 1y agoI'm so surprised there is so much pushback against this.. AWS is extremely expensive. The use cases for setting up your system or service entirely in AWS are more rare than people seem to realise. Maybe I'm just the old man screaming at cloud (no pun intended) but when did people forget how to run a baremetal server ? > We have 730+ days with 99.993% measured availability and we also escaped AWS region wide downtime that happened a week ago. This is a very nice brag. Given they are using their ddos protection ingress via CloudFlare there is that dependancy, but in that case I can 100% agree than DNS and ingress can absolutely be a full time job. Running some microservices and a database absolutely is not. If your teams are constantly monitoring and adjusting them such as scaling, then the problem is the design. Not the hosting. Unless you're a small company serving up billions of heavy requests an hour, I would put money on the bet AWS is overcharging you.
- JCM9 1y agoAs the author points out AWS can provide a few things that you wouldn’t want to try and replicate (like CloudFront) but for most other things you’re very much correct. AWS is ultimately very expensive for what it is. The complicated billing that’s full of surprises also makes cost management a head-banging experience.
- tyingq 1y agoFair, though using AWS solely for CloudFront would mean you should compare to Cloudflare, Akamai, Fastly, etc. I'm not sure if the value prop for it looks so great if you don't include the "integrated with your other AWS stuff" benefit.
- vidarh 1y agoI mean, AWS egress is so expensive that I'd put something else in front of it for anyone who has any decent amount of traffic.
- JCM9 1y agoAgree, CloudFront isn’t super competitive with CDN focused vendors. It’s basically the “well you’re already on AWS so may as well just use this” play.
- seidleroni 1y agoAs someone who works with firmware, it is funny how different our definitions of "bare metal" is.
- embedding-shape 1y agoAs someone who does material science, it's funny how our definition of "bare metal" is so different.
- onionisafruit 1y agoAs someone who listens to loud rock and roll music …
- amluto 1y agoAsk an astronomer what a “metal” is.
- andrewl-hn 1y agoIn similar way I once worked on a financial system, where a COBOL-powered mainframe was referred to as "Backend", and all other systems around it written in C++, Java, .NET, etc. since early 80s - as "Frontend".
- embedding-shape 1y agoHad somewhat similar experience, the first "frontend" I worked on was a sort of proxy server that sat in front of a database basically, meant as a barrier for other applications to communicate via. At one point we called the client side web application "frontend-frontend" as it was the frontend for the frontend.
- pgwhalen 1y agoI don't work in firmware at all, but I'm working next to a team now migrating an application from VMs to K8S, and they refer to the VMs as "bare metal" which I find slightly cringeworthy - but hey, whatever language works to communicate an idea.
- cs702 1y agoIn the early days of cloud service providers, they offered a handful of high-value services, all at great prices, making them cost-competitive with bare metal but much easier. That was then. Things today are different. As cloud service providers have grown to become dominant, they now offer a vast, complicated tangle of services, microservices, control panels, etc., at prices that can spiral out of control if you are not constantly on top of them, making bare metal cheaper for many use cases.
- embedding-shape 1y ago> they offered a handful of high-value services, all at great prices, making them cost-competitive with bare metal but much easier That was never the case for AWS, the point was never "We're cheap" but "We let you scale faster for a premium". I first came across cloud services around 2010-2011 I think, when the company I worked at at the time started growing and we needed something better than shared hosting. AWS was brought up as a "fresh but expensive" alternative, and the CTO managed to convince the management that we needed AWS even if it was expensive, because it'll be a lot easier to tear up/down servers as we need it. Bandwidth costs I think was the most expensive part of the package, at least back then. When I look at what performance per $ you get with AWS et al today, it looks the same, incredibly expensive for the performance you (don't) get. Better off with dedicated instances unless you team is lacking the basic skills of server management, or until the company really grown so it keeps being difficult dealing with the infrastructure, then hire a dedicated person and let them make the calls for what's next.
- everfrustrated 1y agoI'd agree that AWS never sold on being cheaper, but there is one particular way AWS could be cheaper and that is their approach to billing-by-the-unit with no fixed costs or minimum charges. Being able to start small from a $1/mth bill without any fixed cost overheads is incredibly powerful for small startups. If I wanted to store bytes in a DC it would cost $10k/mth by the time I was paying colo/ servers/ disks before I stored my first byte. Sure there wouldn't be any incremental costs for the second byte but thats a steep jump. S3 would have cost me $0.02. Being able to try technology and prove concepts at the product development stage is very powerful and why AWS became not just a vendor but a _technology partner_ for many companies.
- doctorpangloss 1y agoMicrok8s has common, catastrophic performance bugs. There are also catastrophic problems with microk8s Ceph addons. So is this post true? Microk8s, for people who know stuff, is a canary for clusters / applications that don’t really work.
- ndhandala 1y agoWe havent found those bugs in our cluster, but we're also moving to Talos (but for diff reasons)
- acejam 1y agoSource? Links?
- doctorpangloss 1y agohttps://www.google.com/search?q=microk8s+dqlite+site%3Agithub.com https://www.google.com/search?q=microk8s+dqlite+site%3Agithu... Click on the various catastrophic issues. Observe how many are closed with no resolution. Canonical is great but Microk8s is not.
- jammo 1y agoEquinix Metal is now EOL, so worth bearing that in mind..
- darkwater 1y agoThe core of this success is this, IMO: > Our workload is 24/7 steady. We were already at >90% reservation coverage; there was no idle burst capacity to “right size” away. If we had the kind of bursty compute profile many commenters referenced, the choice would be different. Which TBH applies to many, many places, even if they are not aware of it.
- marcinzm 1y agoI'd say the core of their success is running everything in a single rack in a single datacenter at first (for months? a year?) and getting lucky. Life is simple when you don't need the costs and effort of reliability upfront.
- darkwater 1y agoThey mention having a second half-rack in a different DC. In any case, not everyone need five nines, and usually it's just much easier to bring down a platform due to some bug in your own software rather that the core infrastructure going down at a rack level.
- sceptic123 1y agoThe point is valid, they mention adding that, so at one point they didn't have that. They're also only storing monitoring & observability data, that's never going to be mission critical for their customers. It's probably the main reason why they were able to get away with this and why their application does not need scalability. I see they themselves are only offering two 9s of uptime.
- Hardwired8976 1y agoThey mentioned having a backup AWS cluster that would spin up when something happens.
- jdsully 1y agoEven if you have that you'll find AWS is "out of stock" and wants you to create reservations that essentially cost the same as just having the machine 24/7.
- cornfieldlabs 1y ago> Equinix Metal got the closest, but bare metal on-demand still carried a 25-30% premium over our CapEx plan. Their global footprint is tempting; we may still use them for short-lived expansion. > The Equinix Metal service will be sunset on June 30, 2026. https://docs.equinix.com/metal/ https://docs.equinix.com/metal/
- mythz 1y agoSeveral years off AWS, the only thing I still prefer AWS for is SES, otherwise Cloudflare has the more cost effective managed services. For everything else we use Hetzner US Cloud VMs for hosting all App Servers and Server Software. Our .NET Apps are still deployed as Docker Compose Apps which we use GitHub Actions and Kamal [1] to deploy. Most Apps use SQLite + Litestream with real-time replication to R2, but have switched to a local PostgreSQL for our Latest App with regular backups to R2. Thanks to AI that can walk you through any hurdle and create whatever deployment, backup and automation scripts you need, it's never been easier to self-host. [1] https://docs.servicestack.net/kamal-deploy https://docs.servicestack.net/kamal-deploy
- sondr3 1y ago> Cloud makes sense when elasticity matters; bare metal wins when baseload dominates. This really is the crux of the matter in my opinion, at least for applications (databases and so on is in my opinion more nuanced). I've only worked at one place where using cloud functions made sense (keeping it somewhat vague here): data ingestion from stations that could be EXTREMELY bursty. Usually we got data from the stations at roughly midnight every day, nothing a regular server couldn't handle, but occasionally a station would come back online after weeks or new stations got connected etc which produced incredible load for a very short amount of time when we fetched, parsed and handled each packet. Instead of queuing things for ages we could instead just horizontally scale it out to handle the pressure.
- marcinzm 1y agoThey were running for a long time (months? over a year?) on a single rack in a single datacenter. Eventually they scaled out but the word is eventually. I think that summarizes both sides of this debate in a nutshell. You can move off of AWS but unless you invest a lot you will take on increased risk. Maybe you'll get lucky and your one rack won't burn down. Maybe you won't. They did get lucky.
- athrowaway3z 1y agoFrom the story, they seem to have kept the option to fallback on AWS.
- gizzlon 1y agoHm.. I wonder what the risk of a rack going offline is? Maybe 5% in a given year? Less? More? Compared to all the other things that can and will go wrong, this risk seems pretty small, but I have no data to back that up.
- shakow 1y ago> Maybe you'll get lucky and your one rack won't burn down Given the rates of fires in DCs, you'd rather need to be quite unlucky for it to happen to you.
- cornfieldlabs 1y agoManaged DB costs a lot. Is there a simple safe setup that we can run on an Ubuntu server? We self-host the Postgres db with frequent backups to s3 but just in case the site takes off, we need an affordable reliable solution. Does anyone here run their own db servers? Any advise? Backups, security, upgrades etc
- lofties 1y agoI love the argument that Managed DBs cost a lot, but they're supposedly safer. Meanwhile people can't figure out the IAM permission models so they give the entire world access with root:root.
- ndhandala 1y agoIf you're running k8s cluster. Check out cloudnative pg. That thing is a beast.
- cornfieldlabs 1y agoWe have hosted on everything on a tiny Hetzner. The site barely has any users apart from our friends:) :( Info noted
- vpShane 1y agoWorth checking out the different server hosts. You can get a cheap OVH server with 64GB of RAM, 4-6cores with 2TB of disk space from OVH for $30, better servers for $70 with 1gbps - 2gbps bandwidth. Setting up a DB isn't hard, using an LLM to ask questions will guide you to the right places. I'm always talking with Gemini because I switched from Ubuntu to Fedora 42 server and things are slightly different here and there. But, different server hosts offer DB-ready OS's so all you have to do is load the OS on the server and you'll be ready to go. The joy of Linux is getting everything _just right_ and so much _just right_ that you can launch a second server and set it up that way _just right_ within minutes.
- film42 1y agoMaybe look at R2 or Wasabi instead of S3. That would cut your storage bill by 3x and take your cloud network bill to zero. IMO self-managing DBs always sucks no matter what you do.
- blindriver 1y agoHave they done a complete failover to their second data center? It wasn’t clear how committed of a failover it was during the tests.
- aeve890 1y ago>We're now moving to Talos. We PXE boot with Tinkerbell, image with Talos, manage configs through Flux and Terraform, and run conformance suites before each Kubernetes upgrade. Gee, how hard is to find SE experts in that particular combination of available ops tools? While in AWS every AWS certified engineer would speak the same language, the DIY approach surely suffers from the lack of "one way" to do things. Change Flux with Argo for example (assuming the post is talking about that Flex and no another tool with the same name), and you have a almost completely different gitops workflow. How do they manage to settle with a specific set of tools?
- zppln 1y agoIf you're that much of a slave to your tool chain you don't get to call yourself an engineer.
- film42 1y agoOr you have PTSD after 10 years of being on-call 24/7 for your company's stack. I've built my next chapter around offloading the pager. Worth every penny.
- 63stack 1y agoArgocd and flux are "almost completely different"? The last time I looked was about a year ago, and there seemed to be only minor differences. What are the major differences?
- amluto 1y agoI would not want to hire an engineer who claimed to be proficient with any cloud Kubernetes stack but couldn’t learn Talos in a week.
- Ragnarork 1y ago> Gee, how hard is to find SE experts in that particular combination of available ops tools? You find expert in Ops, not in tools. People that know the fundamentals, not just the buttons to push in "certain situations" without knowing what's really going on under the hood.
- 3626351828272 1y ago[dead]
- ecshafer 1y agoAWS is extremely expensive, and I think I have to agree with DHH's assessment that many developers are afraid of computers. AWS is taking advantage of that fear of actually just setting up linux and configuring a computer. However to steelman AWS use. Many businesses are STILL running mainframes. Many run terrible setups like Access as a production database. In 2025 there are large companies with no CICD platforms or IAC, and some companies where even VC is still a new concept or a dark art. So not every company is in the position to actually hire competent system administrators and system engineers to set up some bare metal machines and configure Ceph, much less Hadoop or Kubernetes. So AWS lets these companies just buy this capabilities while forcing the software stack to modernize.
- faxmeyourcode 1y agoI worked at a company like this, I was an intern with wide eyes seeing the migration to git via bitbucket in the year ... 2018? What a sight to see. That company had its own data center, tape archives, etc. It had been running largely the same way continuously since the 90s. When I left for a better job, the company had split into two camps. The old curmudgeonly on-prem activists and the over-optimistic cloud native AWS/GCP certified evangelist with no real experience in the cloud (because they worked at a company with no cloud presence). I'm humble enough to admit that I was part of the second camp and I didn't know shit, I was cargo culting. This migration is still not complete as far as I'm aware. Hopefully the teams that resisted this long and never left for the cloud get to settle in for another decade of on-prem superiority lol.
- ecshafer 1y agoI was a at a company that was doing their SVN/Jenkins migration to Git/Bitbucket/Bamboo around 2016/2018. But they were using source control and a build system already, so you have to hand it to them. But I have an associate that was at one of the large health insurance companies in 2024, complaining that he couldn't get them to use git and stop deploying via FTP to a server. There is danger with being too much on the cargo cult side, but also danger with being too resistant to change. I don't know how you can look at source control, a CICD pipeline, artifacts, IaC, and say "This looks like a bad idea".
- iLoveOncall 1y agoThis is a completely meaningless article if they don't provide information about their technical stack, which AWS services they used to use, what TPS they are hitting, what storage size they're using, etc. The story will be different for every business because every business has different needs. Given the answer to "How much did migration and ongoing ops really cost?" it seems like they had an incredibly simple infrastructure on AWS, and it was really easy to move out. If you use a wider-range of services the cost savings are much more likely to cancel themselves.
- globular-toast 1y agoTFA begins with a link to the original article with those details.
- iLoveOncall 1y agoIf you called "We used EKS" details, then yeah they provide those details. Assuming this is indeed all they used, this was admittedly nonsense, they were essentially using cloud-based bare-metal.
- tuhgdetzhh 1y agoQuite recently I made a TCO analysis between AWS and bare metal Hetzner including salary. https://beuke.org/hetzner-aws/ https://beuke.org/hetzner-aws/
- TYPE_FASTER 1y ago> It depends on your workload. Very much this. Small team in a large company who has an enterprise agreement (discount) with a cloud provider? The cloud can be very empowering, in that teams who own their infra in the cloud can make changes that benefit the product in a fraction of the time it would take to work those changes through the org on prem. This depends on having a team that has enough of an understanding of database, network and systems administration to own their infrastructure. If you have more than one team like this, it also pays to have a central cloud enablement team who provides common config and controls to make sure teams have room to work without accidentally overrunning a budget or creating a potential security vulnerability. Startup who wants to be able to scale? You can start in the cloud without tying yourself to the cloud or a provider if you are really careful. Or, at least design your system architecture in such a way that you can migrate in the future if/when it makes sense.
- mr_toad 1y agoThis is a tech company and it’s adjacent to their core competency. Most companies wouldn’t know MicroK8s from a brand of cereal, they’d only create a mess if they tried this themselves.
- gizzlon 1y agoSure, but they also create a mess in AWS
- stuff4ben 1y agoNever heard of Talos before now. That looks pretty cool and I might start playing with that on my home lab. Can't use it at work for reasons, but good to keep on top of tech (even if I am a little behind)
- globular-toast 1y agoThis dude did a complete walkthrough setting up a Talos cluster on bare metal: https://datavirke.dk/posts/bare-metal-kubernetes-part-1-talos-on-hetzner/ https://datavirke.dk/posts/bare-metal-kubernetes-part-1-talo... It's a nice read. I have my own Talos cluster running in my homelab now for over a year with similar stuff (but no Ceph).
- deleted 1y ago[deleted]
- roschdal 1y agoBare metal is the best metal.
- Aldipower 1y agoNever ever. True metal it is!
- ksec 1y agoMany other points. When the Cloud Started, they offered great value in adjacent product and services. Scaling was painful, getting bare metal hardware have long lead time, provisioning takes time. DC was not of as high quality, Network wasn't as redundant. A lot of these today are much less of an issue. In 2010 you could only get 64 Core Xeon CPU coming in 8 Sockets, or maximum or 8 Core per socket. And that is ignoring NUMA issues. Today you could get 256 Core per socket that is at least twice as fast per core. What used to be 64 Server could now be fitted into 1. And by 2030, it would be closer to 100 to 1 ratio. Not to mention Software on Server has gotten a lot faster compared to 2010. PHP, Python, Ruby, Java, ASP or even Perl. If we added up everything I wouldn't be surprised we are 200 or 300 to 1 ratio compared to 2010. I am pretty sure there is some version of Oxide in the pipeline that will catch up to latest Zen CPU Core. If a server isn't enough, a few Oxide Rack should fit 99% of Internet companies usage.
- submeta 1y agoThere is so much hidden cost in maintaining your own bare metal infrastructure. I am always astounded by how people overlook the massive opportunity cost involved in not only setting up, securing, and maintaining your bare metal infrastructure, but also make it state of the art, including best practices, making sure you have required uptime, monitoring and intervening if necessary. - I work in a highly regulated market with 700 coworkers, our IT maintains an endless amount of VMs. And you cannot imagine how much more work they have to do compared to a setup where you spin up services in AWS or Azure. And destroy it when you don’t need it. No updates, no patches. No misconfiguration. Not every company uses automation either (chef, ansible and whatnot)
- saxenaabhi 1y agoI agree, I have a restaurant POS system and I think self-hosting would easily kill the product velocity, and if we screw up bad, even the company. However, I do get the point about cost-premium and more importantly vendor-risk that's paid when using managed services. We are hosted on cloudflare workers which is very cheap, but to mitigate the vendor risk we have also setup up replicas of our api servers on bunny.net and render.com.
- pingoo101010 1y agoMany startups and companies couldn't exist if there was only AWS (or GCP / Azure) due to how much they overcharge. For example, we couldn't offer free GeoIP downloads[0] if we were charged the outrageous $0.09 / GB, and the same is true for companies serving AI models or game assets. But what makes me almost sick is how slow is the cloud. From network-attached disks to overcrowded CPUs, everything is so slooooow. My experience is that the cloud is a good thing between 0-10,000 $ / month. But you should seriously consider renting bare-metal servers or owning your own after that. You can "over-provision" as much as you want when you get 10-20x (real numbers) the performance for 25% of the price. [0] https://downloads.pingoo.io https://downloads.pingoo.io
- hedora 1y agoI’ve seen cloud slowness create weird Stockholm syndrome effects, especially around disk latency. It always makes sense to compare to back of the envelope bare metal numbers before rearchitecting your stack to work around some dumb cloud performance issue.
- ed_mercer 1y agoTalos is great until it's not. We ran into Ceph IO speed bottlenecks and found it was impossible to debug ("talosctl cgroups —preset=io" is a mess) because the devs didn't want to add an SSH escape hatch into their black box OS. Our Talos nodes would also randomly become unhealthy and you have no way of knowing why. Switched to PXE booted Alpine linux with vanille k8s, and we had a much more stable experience with no surprises, and the ability to SSH whenever we want has been hugely helpful.
- dev_l1x_be 1y ago> AWS is extremely expensive. I really like how people throw around these baseless accusations. S3 is one of the cheapest storage solutions ever created. The last 10 years I have migrated roughly 10-20PB worth of data to AWS S3 and it resulted in significant cost saving every single time. If you do not know how to use cloud computing than yes, AWS can be really expensive.
- Aurornis 1y agoThe implicit claims are more misleading, in my opinion: The claim that self-hosting is free or nearly free in terms of time and engineering brain drain. The real cost of self-hosting, in my direct experience with multiple startup teams trying it, is the endless small tasks, decisions, debates, and little changes that add up over time to more overhead than anyone would have expected. Everyone thinks it’s going to be as simple as having the colo put the boxes in the rack and then doing some SSH stuff, then you’re free of those AWS bills. In my experience it’s a Pandora’s box of tiny little tasks, decisions, debates, and “one more thing” small changes and overhauls that add up to a drain on the team after the honeymoon period is over. If you’re a stable business with engineers sitting idle that could be the right choice. For most startups who just need to get a product out there and get customers, pulling limited headcount away from the core product to save pennies (relatively speaking) on a potential AWS bill can be a trap.
- marcosdumay 1y ago> The claim that self-hosting is free or nearly free in terms of time and engineering brain drain. Free? No, it's not free. It only costs less engineering time than AWS.
- rafabulsing 11mo ago> The implicit claims are more misleading, in my opinion: The claim that self-hosting is free or nearly free in terms of time and engineering brain drain. Not only is that not an implicit claim in the post, they explicitly say the that it's not free, it's actually just around the same amount of time they used to spend with AWS: > Total toil is ~14 engineer-hours/month, including prep. The AWS era had us spending similar time but on different work: chasing cost anomalies, expanding Security Hub exceptions, and mapping breaking changes in managed services. The toil moved; it did not multiply. As for the following: > If you’re a stable business with engineers sitting idle that could be the right choice. For most startups who just need to get a product out there and get customers, pulling limited headcount away from the core product to save pennies (relatively speaking) on a potential AWS bill can be a trap. You're just agreeing with them: > Cloud-first was the right call for our first five years. Bare metal became the right call once our compute footprint, data gravity, and independence requirements stabilised.
- spprashant 1y agoThe thing I find counter intuitive about AWS and hyper-scalers in general is, they make so much sense when you are starting out a new project. A few VMs, some gigs of data storage, you are off to the races in a day or two. As soon as you start talking about any kind of serious data storage and data transfer the costs start piling up like crazy. Like in my mind, the cost curve should flatten out over time. But that just doesn't seem to be the reality.
- shadowgovt 1y agoSounds like they did the right thing for their business model. I think as AWS grows and changes the curve of the target audience is changing too. The value proposition is "You can get Cloud service without having a dedicated Cloud team," but there are caveats: - AWS is complicated enough that you will still need a team to integrate against it. The abstractions are not free and the ones that are leaky will bite you without dedicated systems engineers to specialize in making it work with your company's goals. - For small companies with little compute need, AWS is a good option. Beyond a certain scale... It is worth noting that big companies build their own datacenters, they don't rely on someone else's Cloud. Amazon, Google, and Microsoft don't run on each other. - Recently, the cost model has likely changed if a company pokes their head up and runs the numbers, there's, uh, quite a few engineers with deep knowledge of how to build a scalable cloud infrastructure available to hire now for some reason. In fact, a savvy company keeping its ear to the ground can probably snap up some high-tier talent very soon (https://www.reuters.com/business/world-at-work/amazon-targets-many-30000-corporate-job-cuts-sources-say-2025-10-27/ https://www.reuters.com/business/world-at-work/amazon-target...). It really depends on where your company's risk and cost models are. Running on someone else's cloud just isn't the only option.
- yanslookup 1y agoFD: I work at Amazon, I also started my career in a time where I had to submit paper requests for servers that had turn around times measured in months. I just don't see it. Given the nature of the services they offer it's just too risky not to use as much managed stuff with SLAs as possible. k8s alone is a very complicated control plane + a freaking database that is hard to keep happy if it's not completely static. In a prior life I went very deep on k8s, including self managing clusters and it's just too fragile, I literally had to contribute patches to etcd and I'm not a db engineer. I kept reading the post and seeing future failure point after future failure point. The other aspect is there doesn't seem to be an honest assessment of the tradeoffs. It's all peaches and cream, no downsides, no tradeoffs, no risk assessment etc.
- AndroTux 1y agoManaging a complex environment is hard, no matter whether that’s deployed on AWS or on prem. You always need skilled workers. On one platform you need k8s experts. On the other platform you need AWS experts. Let’s not pretend like AWS is a simple one-click fire and forget solution. And let’s be very real here: if your cloud service goes down for a few hours because you screwed something up, or because AWS deployed some bad DNS rules again, the world moves on. At the end of the day, nobody gives a shit.
- yanslookup 1y agoMaybe I've drank the koolaid but I've done both a lot of systems level work and AWS work (I don't actually use any AWS stuff in my role here interestingly) and I think for a business that needs a handful of hosts in 2 AZs I can't imagine the ROI and risk profile being better to self host. AWS truly does let you focus on your business logic and abstracts a TON of undifferentiated work and well beyond the low hanging fruit of system updates and load balancing. I guess put another way, providing a SaaS you need to have an SLA, those SLAs flow from SLO and SLIs and ultimately a risk profile of your hw and sw. The risk of a bad HBA alone probably means a day of downtime if you don't do things perfectly. AWS has bad HBAs, CPUs, memory, disks etc all day long every day and it's not even a blip for customers, never mind downtime. And if you don't model bad HBAs in your SLAs then your board is going to be pissed when that outage inevitably happens. Now if you don't have SLAs and you like sysops, networkops, clusterops, dbops work then sure, YOLO.
- thelastgallon 1y agoThese are the features that AWS provides (1) Massive expansion of budget (100 - 1000x) to support empire building. Instead of one minimum-wage sysadmin with 2 high-availability, maxed-out servers for 20K - 40K (and 4-hour response time from Dell/HPE), you can have 100M multi-cloud Kubernetes + Lambda + a mix-and-match of various locked-in cloud services (DB, etc.). And you can have a large army of SRE/DevOps. You get power and influence as a VP of Cloud this and that and 300 - 1000 people reporting to you. (2) OpEx instead of CapEx (3) All leaders are completely clueless about hiring the right people in tech. They hire their incompetent buddies who hire their cronies. Data centers can run at scale with 5-10 good people. However, they hire 3000 horrible, incompetent, and toxic people, and they build lots of paperwork, bureaucracy, and approvals around it. Before AWS, it was VMware's internal cloud that ran most companies. Getting bare metal or a VM will take months to years, and many, many meetings and escalations. With AWS, here is my credit card, pls gimme 2 Vms is the biggest feature.
- torginus 1y agoThe problem with those 5 people, is you can't hire a 6th - your stack is custom and probably even if you find the guy, he'll need months of ramp-up. In contrast, you could throw a stone into a bush and hit an AWS guy.
- rikafurude21 1y agoIf your 6th needs months to understand how the basic blocks in your system are arranged then he might not be one of the "good" guys
- torginus 1y agoNot really a hardcore infra guy, but on the coding side, I know companies with products that have codebases in the multi million LoC range written over decades, one of my friends interned there and told me they didn't even let him work on the core product for months, they put him on some custom testing framework they had for it, just so he could get familiar enough with the core code to be able to contribute meaningfully. He told me that before they started doing that, there were incidents like teams writing entire modules they didn't know already existed - now there were 2 pieces of code doing basically the same thing, that were just incompatible enough to not be possible to merge them.
- rglover 1y agoI really dislike how this industry oscillates between various states of epiphany that things that are overcomplicated and expensive are overcomplicated and expensive. As an industry, we must look like utter clowns to the world. It's really sad that saying "own or control your own servers" seems to be a sword in the stone moment for far more people than it should. Things that used to be a "duh" are now a "wow" and it's deeply unsettling to watch.
- dimitrios1 1y agoOne thing I can say definitively, as someone who is definitely not an AI zealot (more of an AI pragmatist): GPT language models have reduced the barrier of running your own bare metal server. AWS salesfolk have long often used the boogeyman of the costs (opportunity, actual, maintenance) of running your own server as the reason you should pick AWS (not realizing you are trading one set of boogeymen for another), but AI has reduced a lot of that burden.
- sema4hacker 1y agoAnycast, Argo Rollouts, Aurora Serverless, AWS, BGP, Ceph, ClickHouse, Cloudflare, CloudFront, DWDM, Flux, Frankfurt, Glacier, Helm, Kinesis, Kubernetes, Metabase, MicroK8s, NVMe, OneUptime, OpenTelemetry Collector, Paris, Postgres, Posthog, PXE, Redis, Step Functions, Supermicro, Talos, Terraform, Tinkerbell, VM's. I wish you started out by telling me how many customers you have to serve, how many transactions they generate, how much I/O there is.
- rossdavidh 1y agoI had a problem figuring out why the place I was working wanted to move from in-house to AWS; their workload was easily handled by a few servers, they had no big bursts of traffic, and they didn't need any of the specialized features of AWS. Eventually, I realized that it was because the devs wanted to put "AWS" on their resumes. I wondered how long it would take management to catch on that they were being used as a place to spruce up your resume before moving on to catch bigger fish. But not long after, I realized that the management was doing the same thing. "Led a team migration to AWS" looked good on their resume, also, and they also intended to move on/up. Shortly after I left, the place got bought and the building it was in is empty now. I wonder, now that Amazon is having layoffs and Big Tech generally is not as many people's target employer, will "migrated off of AWS to in-house servers" be what devs (and management) want on their resume?
- whstl 1y agoDevs wanting to put AWS on their resume push for it, then the next wave you hire only knows AWS. And then discussions on how to move forward are held between people that only know AWS and people who want to use other stuff, but only one side is transparent about it.
- ahel 1y agowith "dev wanting X" nothing happens. "leadership deciding X" then it needs to get done.
- rossdavidh 1y agoThat was what I thought, but when the middle management also wants it, then it can become the 'obvious choice', a la 'nobody ever got fired for choosing IBM'. It seems that middle management + devs can make it seem inevitable to the people above them, especially if those people are non-tech.
- hedora 1y agoReason to use AWS from the article: > You do not have the appetite to build a platform team comfortable with Kubernetes, Ceph, observability, and incident response. Has work been using AWS wrong? Other than Ceph, all those things add up to onerous half time jobs for rotating software engineers. Before gp3 came out, working around EBS price/performance terribleness was also on the list.
- electroly 1y agoI put our company onto a hybrid AWS-colocation setup to attempt to get the best of both worlds. We have cheap fiddly/bursty things and expensive stable things and nothing in between. Obviously, put the fiddly/bursty things in AWS and put the stable things in colocation. Direct Connect keeps latency and egress costs down; we are 1 millisecond away from us-east-1 and for egress we pay 2¢/GB instead of the regular 9¢/GB. The database is on the colo side so database-to-AWS reads are all free ingress instead of egress, and database-to-server traffic on the colo side doesn't transit to AWS at all. The savings on the HA pair of SQL Server instances is shocking and pays for the entire colo setup, and then some. I'm surprised hybrids are not more common. We are able to manage it with our existing (small) staff, and in absolute terms we don't spend much time on it--that was the point of putting the fiddly stuff in AWS. The biggest downside I see? We had to sign a 3 year contract with the colocation facility up front, and any time we want to change something they want a new commitment. On AWS you don't commit to spending until after you've got it working, and even then it's your choice.
- jcalvinowens 1y agoI have seen multiple startups paying thousands of dollars a month in AWS bills to run a tiny service which could trivially run on an $800 desktop on a residential internet connection. It's absolutely tragic.
- hedora 1y agoThat’s like $24K a year. Assuming they have working failover and business continuity plans, it’s actually a really good deal (vs having a 10-20% time employee deal with it).
- deleted 1y ago[deleted]
- whstl 1y agoAWS doesn't get magically expensive just because you put your website there. You don't get to an overcomplicated AWS madness without having a few engineers already pushing complexity. And an overcomplicated setup also means it needs maintenance. There are no personnel savings there.
- hedora 1y agoFor one VM, EBS with backups gives you business continuity. You could get manual failover with a single writer replicated managed Postgres setup and a warm VM. That’s on the order of a thousand a month for a medium workload. It’s probably a 10x markup vs buying the servers, but it doesn’t matter if it saves an employee.
- whstl 1y agoIt doesn’t save employees. Over-complicated infrastructure doesn’t magically appear out of nowhere. Someone has to setup and maintain. It’s expensive.
- kaydub 1y agoThen they architected and built it out wrong. It's pretty simple to run low cost services in AWS. If you're small enough that an $800 desktop on a home internet connection will handle it, surely you could run it completely serverless for much less. I'm surprised how many people I see wanting to go on-prem vs AWS/public cloud. Feels penny smart and pound foolish to me. Lots of people too deep into the technical side of things that they don't even understand the business side.
- debarshri 1y agoRecently i learned that orgs these days want to show software and infrastructure spend as capex as they can shown it as depreciating asset for tax purposes. I understand that with AWS you cannot do that as it is often seem as opex. I guess thats a good enough motivation to move out of AWS at scale.
- kyledrake 1y agoThe article mentions Equinix Metal but if you look it up they are shutting down the service https://docs.equinix.com/metal/hardware/standard-servers https://docs.equinix.com/metal/hardware/standard-servers Doesn't make me want to be a Equinix customer when they just randomly shut down critical hosting services. I'm pretty sure that it's just the post-merger name for Packet which was an incredible provider that even had BYO IP with an anycast community. Really a shame that it went away, it was a solid alternative to both AWS and bare metal and prices were pretty good. There's a missing middle between ultra expensive/weird cloud and cheap junk servers that I would really love to see get filled.
- dilyevsky 1y agoFwiw equinix metal was an acquisition (Packet). Seems like it didnt go too well
- jameson 1y agoCurious to know how's the development experience been post-migration? Was there additional friction due to lack of tooling in on-prem that would otherwise available in the cloud env for example?
- carlgreene 1y agoOk so this may be a dumb question...but now do you handle ISP outages due to storms and stuff with on prem solutions? I'd imagine large datacenters have much more sophisticated and reliable internet connections than say an Xfinity business customer, but maybe that's wrong.
- neuronflux 1y agoMuch more sophisticated and reliable than Xfinity. Good datacenters have redundant and physically separated power and communication from different providers. Also, in case something catastrophic happens at one datacenter, the author mentions they are peered to another datacenter in a different country, as another layer of redundancy. Cloudflare handles their ingress, so such a catastrophic event wouldn't likely to be noticed by their customers.
- flufluflufluffy 1y agoRight! I can’t believe they decided to ditch the OS entirely and maintained availability like that!
- yearolinuxdsktp 1y agoRunning EKS on AWS was their problem. If they didn't run EKS on AWS, they would've had a considerably simpler setup running Amazon Linux, not having to upgrade Kubernetes every 3 quarters, managing network security using security groups instead of having open internal networking, and running in a single AZ would've eliminated intra-AZ costs. In large data centers like us-east-1, an individual AZ is actually internally striped for extra redundancy, and you are much more likely to experience regional downtime than single AZ downtime, especially if you have a stable workload and do not rely on tech beyond rock-solid basics (EC2, VPC, ELB, S3, EBS). If you're willing to operate a single bare metal rack in a DC, you should be willing to run in a single AWS AZ. I don't know how much time they spend configuring/dealing with Kubernetes, but I bet it's a large chunk of the 24 hour engineer-hours per quarter. But this is not a required expense: "EKS had an extra $1,260/month control-plane fee". Running EKS adds a massive IAM policy maintenance overhead, whereas a non-EKS (EC2 w/ golden AMIs) setup results in drastically simpler IAM policies. NAT gateways are ~$50 a month, plus data transfer. Setting up a gateway VPC endpoint to S3 will avoid having to pay transfer charges to S3. They were at 90% reservation capacity, so they should be using reservations for greater savings and in fact, running stable workloads with reservations is something that AWS excels at. Reservation means that you will be able to terminate and re-launch instances even when there's a spike in demand from other users--your instance capacity is guaranteed. Running the basics on VMs also effectively avoids vendor lock-in. Every cloud provider supports VMs with a RedHat clone, VPCs, load balancing, networked storage, access controls, object storage and a fixed size fleet with auto-relaunch on instance failure. With a consistent workload, they would have very likely escaped the downtime from AWS a week ago as well, because, as per AWS, "existing EC2 instances that had been launched prior to the start of the event remained healthy and did not experience any impact for the duration of the event". With Terraform and automation for building launchable images, you can stand up a cluster quickly in any region with secure networking, including in a separate AWS account, in the same region, for the sake of testing. With AWS, you can set up automatic EBS backups of all your data to snapshots trivially, and even send them to a 3rd locked-down account, so they can't be accidentally wiped.
- ZebusJesus 1y agoThank you for the share this is really good information for making expensive decisions!
- nemothekid 1y ago>We now save over $1.2M / yr and we expect this to grow, as we grow as a business. Am I just naive? How is a uptime SaaS product saving over a million year on managed colo vs AWS? Was every API route in it's own EC2 instance? AWS is expensive sure, but over a million dollars a year? For this product specifically?. I got some clarification from their earlier posts and it looks like they were intentionally avoiding any AWS platform features: >Our goal was to avoid reliance on AWS or any proprietary cloud technology. >When we were utilizing AWS, our setup consisted of a 28-node managed Kubernetes cluster. Each of these nodes was an m7a EC2 instance. With block storage and network fees included, our monthly bills amounted to $38,000+. This brought our annual expenditure to over $456,000+. I just think if you are going to deploy on AWS, then treat it AWS like managed-colo, then your bill is going to be high. I understand how that seems unfair, but AWS isn't really in the business of selling virtual machines. If you sit down and ask yourself how you got here, it just seems like you committed yourself to wasting money. If I knew I just needed some linux boxes from the start, there are better choices than AWS.
- nostrebored 1y agoIn almost 100% of cases I've seen this, people are convinced that they are going to go for a multi-cloud situation and using any value added service is "lock-in". The amount of times I've seen the migration scenario play out favorably for people has been absolutely ZERO. Meanwhile, I've seen huge companies successfully complete cloud->cloud migrations in less than a year, as long as they use the value added services of the other cloud.
- pjdesno 1y agoI'm involved in a fairly large academic cloud deployment, sited in a 15MW data center built and shared by a few large universities. There are huge advantages of scale to computer operations in a few areas: - facility: the capital and running cost of a purpose-built datacenter is far cheaper per rack than putting machines in existing office-class buildings, as long as it's a reasonable size - ours is ~1000 racks, but you might get decent scale at a quarter of that. (also one fat network pipe instead of a bunch of slow ones) - purchasing: unlike consumer PCs, low-volume prices for major vendor servers are wildly inflated, and you don't get decent prices until you buy quite a few of them. - operations: people come in integer units, and (assuming your salary ranges are bounded) are only competent in small number of technical areas each. Whether you have one machine or 1000s you need someone who can handle each technology your deployment depends on, from Kubernetes to network ops; multiply 4x for those requiring 24/7 coverage, or accept long response times for off-hours failures. That last one is probably the kicker. To keep salary costs below 50% of your total, assuming US pay rates and 5-year depreciation since machines aren't getting faster as quickly as they used to, you probably need to be running tens of millions of dollars in hardware. Note that a tiny deployment of a few machines in a tech company is an exception, since you have existing technical staff who can run them in their spare time. (and you have other interesting work for them to do, so recruiting and retention isn't the same problem as if their only job was to babysit a micro-deployment) That's why it can be simultaneously true that (a) profit margins on AWS-like services are very high, and (b) AWS is cheaper than running your own machines for a large number of companies.
- kshacker 1y ago> the capital and running cost of a purpose-built datacenter is far cheaper per rack than putting machines in existing office-class buildings, as long as it's a reasonable size - ours is ~1000 racks, but you might get decent scale at a quarter of that. Just want to confirm what I am reading. You are talking about ~1000 racks as the facility size, not what a typical university requires.
- Frannky 1y agoI only use bare metal—super cheap and very easy to switch. No worries about crazy bills or handling the crazy complexity of their systems. So far, so good. When/if problems start, I'll try them
- bhewes 1y agoYes to this keep core base load in your own bare metal systems, use the clouds for what they do best.
- unixhero 1y agoI went bare metal too. Not because of AWS, but because of being frozen out by Hetzner because of a debt of 0.02eur with no way of paying it.
- StratusBen 1y agoCo-Founder and CEO of https://vantage.sh/ https://vantage.sh/ here - I've been pretty impressed by the rate that repatriation is happening off of public cloud. It rarely ever came up and in the last year it's been popping up more and more -- and especially just for getting access to GPU workloads. I thought there would be a greater unbundling to AWS or to cheaper providers but it seems like a good-sized portion of the market is just going back to managing their own hardware.
- Naklin 1y ago> We spent a week of engineers time (and that is the worst case estimate) on the initial migration, spread across SRE, platform, and database owners. I’m sorry but I don’t believe this for one second. And unfortunately that makes me distrust the entirety of the article.
- jiggawatts 1y agoSomething I've noticed about all public clouds is that they promise a unique value proposition, and then entirely fail to deliver it for almost all customers. Specifically, they promise to provide a "small, rapidly growable slice of a very big thing". I.e.: You can create an empty S3 account for $0.00, fill it with a few megabytes of initial data for $0.00001, and then if you suddenly need petabytes then it scales smoothly and beautifully up past any reasonable scale. You get billed for that at an exorbitant rate, but the point is that you can do it without having to rearchitect anything. "This Works Great(tm)" for three specific categories of customers: - Tiny organisations that expect to grow big suddenly: Startups, and maybe the few orgs that have rare or annual events and nothing in between. Think electronic voting systems and the like. - Small international companies that need global presence on the cheap. SaaS vendors, IoT, and a few niche organisations are pretty much the only ones in this category. - Enormous organisations that need large scale but can't be bothered (or can't afford) managing that. There was an AWS talk about a customer that needed about 1 PB of storage which these days is "just" 500 x 20 TB disks, but they needed burst IOPS far in excess of that, on the order of 100,000 disks. The "thin slice" model of the cloud works great for this, because S3 has millions of disks behind it, and each customers' data is spread out over those disks. Everybody else in the "medium" category is sold a bunch of bullshit. Cloud VMs are between 5x and 10x as expensive as the on-prem equivalents, all costs factored in. I've seen the numbers from CIOs and CTOs, they all got told "the cloud is cheaper", and their costs skyrocketed as soon as they went to the cloud. You still need engineers. You still need "deployments". You still need sysops. You still need to update your VMs and their software somehow. Nothing changes really, except suddenly VMware starts looking cheap in comparison. The "You can scale, the cloud is flexible, you can..." marketing is a load of bollocks. First, Pay-as-you-Go pricing is on average 7x on-prem VM pricing. The only way to bring this down to merely 3x the on-prem cost is to LOCK IN the compute using "reservations" of some sort, typically for 3-year periods. This is NOT FLEXIBLE BY DEFINITION! The whole marketing of the cloud revolves around the flexibility, but all of their pricing and cost optimisation revolves around getting customers to lock in spending for years and years. Similarly, Spot-priced anything is so unreliable that it is completely unusable for almost all "enterprise" customers, even for non-production use. Seriously, can you imagine telling a developer that costs $200/hour that they can't do their work because their DEV instance is gone for a day because some other tenant needed it more!? This underlying issue with VM pricing means that any service that is built on top of VMs inherits the same pricing model with lock-in contracts, totally negating the scalability benefits and global presence of the public cloud. If you want S3 in every region with 10 MB each... that's cheap. If you want just one VM in every region... no longer cheap. If you want any VM-based service in every region... also not cheap. Oh... you wanted DEV/TST/UAT/GREEN/BLUE? With high availability? Get the CFO on the line, he'll need to approve your budget! "Just engineer your software to be cloud native!" is what you inevitably hear from apologists. Sure, sure... I'll get right on that. I mean sure, the public cloud vendors failed to do so for like two thirds of their own first-party products, but I'm sure I'll have better luck! Let's see.. my "medium sized org" has... checks notes... about 1,000 unique pieces of software deployed on VMs, of which 800 are CotS vendor products, 600 of which require Windows Server and have a GUI configurator. This will go... smooth. For 90% of the potential customers out there, the big businesses, the enterprises, the universities, governments, and the like... it's just a more expensive data centre that someone else runs for them. The only significant advantages to public cloud vs on-prem I've seen are: - Faster networking, with a well-engineered 100 Gbps or 200 Gbps core in most clouds. This includes Internet uplinks of similar spec. These are rare in private hosting. - Three-way zone redundancy instead of the typical two-way. - Zone redundant services that are just a "checkbox". - Separation of duties where the layer 2 network and hypervisor are managed by a vendor instead of internal staff. (This can be difficult to arrange internally for medium sized orgs needing high security, there aren't enough IT admin staff for true separation.) ... that's about it.
- twodave 1y agoI feel like it just depends on what you’re trying to do. * data driven website with some internal api integration, maybe some client-side application or tooling? Put a server rack in a closet and get a fiber line. * trying to serve the general public in a bursty, not-cacheable way? Probably going to have to carry a lot of machines that usually don’t do much, cloud might make more sense * lots of ingress or DNS rules? A hybrid approach could make sense. In general once you start thinking about scaling data to larger capacities is when you start considering the cloud, because just the storage solution ends up on a long amortization schedule if you need it to be resilient, let alone the servers you’re racking to drive the DB.
- layoric 1y ago> In general once you start thinking about scaling data to larger capacities is when you start considering the cloud What kind of capacities as a rule of thumb would you use? You can fit an awful lot of storage and compute on a single rack, and the cost for large DBs on AWS and others is extremely high, so savings are larger as well.
- twodave 1y agoWell, if you want proper DR you really need an off-site backup, disk failover/recovery, etc. And if you don’t want to manually be maintaining individual drives then you’re looking at one of the big, expensive storage solutions with enterprise grade hardware, and those will easily cost some large multiple more than whatever 2U db server you end up putting in front of it.
- nodesocket 1y ago“ $600/month for NAT gateways” I build my own NAT instances from Debian Trixie with Packer. The configuration is literally a few lines: sudo iptables -t nat -A POSTROUTING -o ens5 -j MASQUERADE sudo iptables -F FORWARD sudo iptables -A FORWARD -i ens5 -m state --state RELATED,ESTABLISHED -j ACCEPT sudo iptables -A FORWARD -o ens5 -j ACCEPT sudo iptables-save | sudo tee /etc/iptables/rules.v4 > /dev/null
- prabhatjha 1y agoThis writeup is very informative. Learned about few OSS frameworks that I was not aware of. Amazing engineering work. Kudos to you all.
- suralind 1y agoLove this write up. But also love the fact that the company pretty much uses the best solutions that they can (that definitely helps with hosting, regardless of cloud/self-host, but is impressive anyway). I’ve been a big Talos fan for couple years, but never worked with a client that would use something different than managed offering from their cloud provider. Similarly, the fact that they are regularly updating Kubernetes, Talos and presumably other things means that these things are a normal flow and thus can optimize and automate. If you’re doing that because AWS is forcing you to upgrade… then it’s stressful.
- stack_framer 1y agoMy question is: How did you get your entire team on board with this decision? My team disagrees on even the most trivial technicalities, so I can't imagine doing something on the scale you're doing (even though I wish we could leave AWS for our own hardware).
- znpy 1y ago> 37% off instances still leaves you paying list price for bandwidth, which was 22% of our AWS bill. This is the most infuriating pain point when using AWS. We have premium support at work and the engineers basically read the documentation back to us, ignoring our complaints about cross-az bandwidth and the utter and complete lack of az-awareness in their services. Example: we had some redises that we wanted to migrate to ElastiCache... All the aws engineers did was reading back the marketing pages and pushing for the serverless offering (where bandwidth was the main cost driver). The reader/writer endpoints are totally az-unaware. Cross-AZ bandwidth for ElastiCache was costing us MORE than ElastiCache itself. In the end we had to engineer AROUND ElastiCache in order to get it working, and working with predictable pricing. "Invent and simplify" my ass...
- babra1 11mo ago[flagged]
- smthglbert7 11mo ago[dead]
- insaneisnotfree 11mo ago“You lean heavily on managed services (Aurora Serverless, Kinesis, Step Functions) where the operational load is the value prop.” Not viable even when your core belongs to AWS. Why? Ask prime video