5 ms·
Having done a bunch of bare metal, I can tell you the calculus isn't really that hard. Bare metal will save you money. Operating bare metal at scale requires
by halbritt 8y ago
Having done a bunch of bare metal, I can tell you the calculus isn't really that hard. Bare metal will save you money.
Operating bare metal at scale requires talent that doesn't exist, not necessarily at an engineering level, but at all levels.
As an example, I worked at a place that had a large bare metal deployment, i.e. >1MW worth of compute. It was woefully inefficient and costly to operate. The product that they offered required network QOS and compute with real time capabilities, neither of which was available from any cloud provider at the time.
One of our executives (formerly a leader in the DC ops org at AWS) left the company to be replaced by another executive by another well-known silicon valley org who then insisted we should migrate everything to the cloud.
I showed him the relatively easy math that efficiently utilized bare metal was way less costly and that the aforementioned QOS and RT requirements would be a deal breaker anyway. He failed to fully grok this and remained insistent. When I quite, he seemed surprised. After the fact, I discovered that they'd made a deal with IBM to move everything into their cloud. A year later it was an utter failure and they abandoned the project.
There are lots of folks in the valley with lots of experience on their resumes that suggests that they should be capable of understanding these kinds of things that simply don't. Lacking that understanding leads to poor decision-making, which leads to failure, which leads to risk-aversion, which leads to everyone believing that it must be cheaper in the cloud.
Or so goes the old adage, "nobody ever got fired for buying IBM."
- throwawaymath 8y agoFeel free to ignore this, but would you shoot me an email? I'm working on a project that uses bare metal and would like to pick your brain. EDIT: To whoever downvoted this, the commenter hasn't listed an email address, or I would have reached out directly. This is an honest attempt at communication that doesn't require someone to break anonymity.
- asdfafd2sdf 8y agoYou were probably downvoted for using the term 'pick your brain'. I am sure they want their brain 'picked'.
- munchbunny 8y agoSure, it's an overused term, but the ask is pretty clear. "I'm doing something and it looks like you've done it before, can I get advice?"
- BenMorganIO 8y agoAgreed. I think it's healthy to ask these things and I believe it leads to a positive and supportive community.
- halbritt 8y agoAgreed. Though the question I received was somewhat nonsensical, which was to be expected.
- halbritt 8y agomy username at gmail dot com
- python300 8y agoThe description here seems to be more of the compute and storage. How about the boat load of services that are offered with AWS. Plugging and playing with services maintained by AWS makes it easier for companies to focus on their product logic. The major expense is actually engineering.
- halbritt 8y agoOutside of S3 I can't think of any services that AWS offers that are worth a damn.
- bradstewart 8y agoRoute53, RDS, DynamoDB, SNS. To name a few.
- skolsuper 8y agoVPC?
- halbritt 8y agoVPC simply is what's there. Google's offerings in that regard are superior. Any SDN on private cloud would be be way, way better.
- halbritt 8y agoRoute53 is okay. Has an API that is incompatible with BIND. There are many other, better providers. Still, it's cheap, so who cares. RDS doesn't really scale without costing a fortune. It buys you HA and backups. Great, but what if you need performance? DynamoDB? It scales in terms of IOPS, but again, it's unaffordable. SNS exists and isn't terrible, but why wouldn't I just run Kafka?
- BurritoAlPastor 8y agoMy forecasting on RDS is that it’s a dead-end product – all the future hotness is going to be in Aurora Serverless. Multi-region active/active Postgres with totally usage-based pricing and totally elastic performance is going to be a game-changer. But if you need bleeding-edge Postgres performance, you hire a DBA, and they probably build something on EC2 or bare metal. ——— As I understand it, RabbitMQ is probably a better point of comparison for SNS/SQS, and Kinesis is the Kafka peer. Regardless, the reason you don’t “just” run Kafka is: you don’t have a team that knows how to tune, deploy, and operate a production Kafka cluster. I learned enough about SNS and SQS to get it running in an afternoon, and I really haven’t needed to think about it since. Kafka (or RabbitMQ, or ActiveMQ, or etc) need instrumentation and monitoring and patching and quorums and capacity planning and etc, and at some scale those are worthwhile, but that scale is MUCH larger than what most Kafka clusters are actually serving. ——— The theme here is: if you have a business requirement for 90th percentile specialized performance, great! Hire domain specialists who can make your systems run at that tier! But for everyone else in the world, when you can get usage-based pricing, elastic resources, and automatic durability and patching... why would you go to the trouble of learning how to deploy and manage a service?
- python300 8y agoFrom a Lyft engineering perspective would rather focus on things like how do I make sense, process, extract the ton of data. How do I focus on customer experience rather than how do I save money in data-center, how do I keep my data-center stack updated and many more
- halbritt 8y agoMy argument is that spending 10% of their revenue on cloud infra affects their unit economics sufficiently that they'll find it difficult to compete. Perfectly willing to admit I'm wrong if and when that time comes. At this point, that's my theory.
- alfalfasprout 8y agoYou can't just eliminate that 10%. Even if going to fully bare-metal lowers costs it takes a lot of time and manpower to make that transition. When that investment can be made in other areas that have much more impact it really doesn't make sense. Bare metal works when your workload is well-defined and understood. Then you can actually put reasonable estimates for what you need and hire/purchase infra accordingly.
- halbritt 8y agoNo disagreement here. The balance here is tricky. Based on public data, it seems that Netflix has ~$16B in revenue against $300m/yr cloud spend. 2% seems much more reasonable to me. I feel like a drive toward efficiency is a worthwhile endeavor for a startup in terms of establishing a competitive advantage.
- eeeeeeeeeeeee 8y agoI've been on both sides and it's not as simple as "bare metal saves you money." It really depends on the company and the type of applications being hosted and where the business is growing (or not growing). An established company with an established workload, especially if it's simple, will probably do better on bare metal, but cloud is popular in the Valley because ideas are still being developed and iterated on heavily where you don't want to be stuck on multi-year hardware leases that might not sync with what that future business looks like. Lyft is probably in that in-between stage, but I still think it's an enormous undertaking to become an infrastructure company over simply being able to hire full-stack developers, which is a lot easier and cheaper. Right now their time is better spent elsewhere. I remember how hard it was to hire senior operations people. There are not many of them, and there are not many of them at the level of being able to deliver something amazing. The ubiquity of the cloud has only made these kind of experts less common. Every place I've worked that did bare metal was always drowning in maintenance instead of working on the next big thing. And no big surprise, our internal infrastructure was nowhere near as high quality or capable as AWS. And most of our developers had experience working directly with cloud providers, without ops people, so we were delivering them a worse experience and slowing them down, and we required more ops people to help them and maintain it and keep everything online. Also, a move to IBM's cloud isn't the greatest example. I had hundreds of bare metal servers in an IBM-owned datacenter and their cloud offering was consistently behind AWS/GCP; if anyone recommended IBM cloud to me I would have laughed at them. It seemed to me that IBM was trying to up-sell on the "cloud" buzz word without actually delivering anything except higher prices, just like how they're now trying to ride the buzz of the blockchain. Dropbox is a good example of a company that took quite a while to move to their own platform, away from AWS (and they still have 10% of their stuff in AWS to this day). Dropbox is basically a storage infrastructure company, unlike Lyft, but it still took them years to invest in the development (and migration) of that custom platform to replace AWS, an investment that not many companies are going to want to gamble on, especially if their primary business is not storage: https://techcrunch.com/2017/09/15/why-dropbox-decided-to-drop-aws-and-build-its-own-infrastructure-and-network/ https://techcrunch.com/2017/09/15/why-dropbox-decided-to-dro... And I think it's telling that Dropbox started on AWS, grew the business on AWS, and moved to a custom platform once their business model was perfected and they wanted to cut costs prior to going public. If Dropbox had started on bare metal from day one, would they have been able to pull it off?
- TheOperator 8y ago>Or so goes the old adage, "nobody ever got fired for buying IBM." This is true but its never really hit me before even though I've already been operating based on the assumption that trusting the cloud is less risky than trusting my own skills.