11 ms·
I cannot overstate the performance improvement of deploying onto bare metal. We typically see a doubling of performance, as well as extremely predictable baseli
by adamcharnock 1y ago
I cannot overstate the performance improvement of deploying onto bare metal. We typically see a doubling of performance, as well as extremely predictable baseline performance.
This is down to several things:
- Latency - having your own local network, rather than sharing some larger datacenter network fabric, gives around of order of magnitude reduced latency
- Caches – right-sizing a deployment for the underlying hardware, and so actually allowing a modern CPU to do its job, makes a huge difference
- Disk IO – Dedicated NVMe access is _fast_.
And with it comes a whole bunch of other benefits:
- Auto-scalers becomes less important, partly because you have 10x the hardware for the same price, partly because everything runs 2x the speed anyway, and partly because you have a fixed pool of hardware. This makes the whole system more stable and easier to reason about.
- No more sweating the S3 costs. Put a 15TB NVMe drive in each server and run your own MinIO/Garage cluster (alongside your other workloads). We're doing about 20GiB/s sustained on a 10 node cluster, 50k API calls per second (on S3 that is $20-$250 _per second_ on API calls!).
- You get the same bill every month.
- UPDATE: more benefits - cheap fast storage, run huge Postgresql instances at minimal cost, less engineering time spend working around hardware limitations and cloud vagaries.
And, if chose to invest in the above, it all costs 10x less than AWS.
Pitch: If you don't want to do this yourself, then we'll do it for you for half the price of AWS (and we'll be your DevOps team too):
https://lithus.eu https://lithus.eu
Email: adam@ above domain
- jnsaff2 1y agoThere is a graph database which does disk IO for database startup, backup and restore as single threaded, sequential, 8kb operations. On EBS it does at most 200MB/s disk IO just because the EBS operation latency even on io2 is about 0.5 ms. Even though the disk can go much faster, disk benchmarks can easily do multi-GB/s on nodes that have enough EBS throughput. On instance local SSD on the same EC2 instance it will happily saturate the whatever instance can do (~2GB/s in my case).
- anonzzzies 1y agoWhat graph db is that?
- jnsaff2 1y agoneo4j
- the_arun 1y agoWhat is the cost of running Neo4j on aws vs using aws Neptune? Related to disk I/o?
- jnsaff2 1y agoTbh, I don't know. For us the switching cost alone would be pretty high. That said ongoing maintenance is pretty high as well.
- zw17 1y agoJust want to chime in. Zhenni, cofounder of PuppyGraph. We created the first graph query engine that can sit on top of your relational databases (think Postgres, Iceberg, Delta lake, etc.), and query your relational data as a graph using Cypher and Gremlin, without any ETL or a separate graphdb needed. It's much more lightweight and easy to spin up. Because we sit on top of column based storage and our compute engine is distributed, we can achieve subsecond query speed across 1 billion nodes. Please check it out!
- inkyoto 1y agoIt is not easily possible to directly compare neo4j and AWS Neptune as the former does not exist as a fully managed service in AWS. neo4j is available through the AWS marketplace, though, but it most assuredly runs on an EC2 instance by neo4j (the company). We run a modest graph workload (relatively small dataset wise but an intense on graph edge wise) on Neptune that costs us slightly under USD 600 per month – that is before the enterprise discount, so in reality we pay USD 450-500 a month. But we use Neptune Serverless that bursts out from time to time, which means that monthly charges are averaged out across the spikes/bursts. The monthly charges are for the serverless configuration of 3-16 NPU's. Disk I/O stats are not available for Neptune, moreso for serverless clusters, and they would not be insightful anyway. The transactions per second rate is what I look at.
- rightbyte 1y agoWhat is old is new again. My employer is so conservative and slow that they are forerunning this Local Cloud Edge Our Basement thing by just not doing anything.
- Aissen 1y agoAs an infrastructure engineer (amongst other things), hard disagree here. I realize you might be joking, but a bit of context here: a big chunk of the success of Cloud in more traditional organizations is the agility that comes with it: (almost) no need to ask permission to anyone, ownership of your resources, etc. There is no reason that baremetal shouldn't provide the same customer-oriented service, at least for the low-level IaaS, give-me-a-VM-now needs. I'd even argue this type of self-service (and accounting!) should be done by any team providing internal software services.
- infogulch 1y agoLike https://oxide.computer/ https://oxide.computer/ ?
- abujazar 1y agoThe permissions and ownership part has little to do with the infrastructure – in fact I've often found it more difficult to get permissions and access to resources in cloud-heavy orgs.
- joshuaissac 1y agoThis could be due to the bureaucratic parts of the company being too slow initially to gain influence over cloud administration, which results in teams and projects that use the cloud being less hindered by bureaucracy. As cloud is more widely adopted, this advantage starts to disappear. However, there are still certain things like automatic scaling where it still holds the advantage (compared to requesting the deployment of additional hardware resources on premises).
- alexchantavy 1y ago
- Thicken2320 1y agoUsing the S3 API is like chopping onions, the more you do it, the faster you start crying.
- scns 1y agoLess to no crying when you use a sharp knive. Japanese chefs say: no wonder you are crying, you squash them.
- Esophagus4 1y agoHaha! My only “yes, but…” is that this: > 50k API calls per second (on S3 that is $20-$250 _per second_ on API calls!). kind of smells like abuse of S3. Without knowing the use case, maybe a different AWS service is a better answer? Not advocating for AWS, just saying that maybe this is the wrong comparison. Though I do want to learn about Hetzner.
- wiether 1y agoThey conveniently provide no detail about the usecase, so it's hard to tell But, yeah, there's certainly a solution to provide better performances for cheaper, using other settings/services on AWS
- adamcharnock 1y agoWe're hoping to write a case study down the road that will give more detail. But the short version is that not all parts of the client's organisation have aligned skills/incentives. So sometimes code is deployed that makes, shall we say, 'atypical use' of the resources available. In those cases, it is great to a) not get a shocking bill, and b) be able to somewhat support this atypical use until it can be remedied.
- wiether 1y agoThank you for the reply I'm honestly quite interested to learn more about the usecase that required those 50k API calls! I've seen a few cases of using S3 for things it was never intended for, but nothing close to this scale
- belter 1y ago> If you don't want to do this yourself, then we'll do it for you for half the price of AWS (and we'll be your DevOps team too You might not realize but you are actually increasing the business case for AWS :-) Also those hardware savings will be eaten away by two days of your hourly bill. I like to look at my project costs across all verticals...
- adamcharnock 1y agoI understand the concern for sure. But we don't bill hourly in that way, as one thing our clients really appreciate is predictable costs. The fixed monthly price already includes engineering time to support your team.
- deleted 1y ago[deleted]
- dlisboa 1y ago> Also those hardware savings will be eaten away by two days of your hourly bill Doubt it. I've personally seen AWS bills in the tens of thousands, he's probably not that costly for a day.
- whstl 1y agoI don't think I have joined a startup that pays less than 20k/month to AWS or any cloud in almost a decade. Biggest recent ones were ~200k and ~100k that we managed to lower to ~80k with a couple months of work (but it went back up again after I left). I fondly remember lowering our Heroku bill from 5k to 2k back in 2016 after a day of work. Management was ecstatic.
- grebc 1y agoThe only real startup I’ve had the privy of looking into financials had 5 figure AWS bills and literally 5 customers. Single digit customers. And the product was sending SMS’s. Startups are clown shows for burning OPM.
- realitysballs 1y agoYa but then you need to pay for a team to maintain network and continually secure and monitor the server and update/patch. The salaries of those professionals , really only make sense for a certain sized organization. I still think small-midsized orgs may be better off in cloud for security / operations cost optimization.
- esskay 1y agoYou still need those same people even if you're running on a bunch of EC2 and RDS instances, they aren't magically 'safer'.
- lnenad 1y agoI mean, by definition yes they are. RDS is locked down by default. Also if you're using ECS/Fargate (so not EC2) as the person writing the article does, it's also pretty much locked down outside of your app manifest definitions. Also your infra management/cost is minimal compared to running k8s and bare metal.
- rightbyte 1y agoIsn't most vulnerabilities in your own server software or configs anyways?
- DisabledVeteran 1y agoThat used to be the case until recently. As much as neither I nor you want to admit it -- the truth is ChatGPT can handle 99% of what you would pay for "a team to maintain network and continually secure and monitor the server and update/patch." Infact, ChatGPT surpasses them as it is all encompassing. Any company now can simply pay for OpenAI's services and save the majority of the money they would have spent on the, "salaries of those professionals." BTW, ChatGPT Pro is only $200 a month ... who do you think they would rather pay?
- tayo42 1y agoYou have a link to some proof that chat gpt is patching servers running databases with no down time or data loss?
- chubot 1y agoDoes anyone have experience with say Linode and Digital Ocean performance versus AWS and GCE? They still use VMs, but as far as I know they have simple reserved instances, not “cloud”-like weather? Is the performance better and more predictable on large VPSes? (edit: I guess a big difference is that VPS can have local NVMe that is persistent, whrereas EC2 local disk is ephemeral? )
- inapis 1y agoNo. DO can be equally noisy but I've always tried their regular instances and not their premium AMD/Intel ones.
- pton_xd 1y agoI can't speak to Linode but in my experience the Digital Ocean VM performance is quite bad compared to bare metal offerings like Hetzner, OVH, etc. It's basically comparable to AWS, only a bit cheaper.
- matt-p 1y agoIt's essentially the same product, but you do get lower disk latency. Best performance is always going to be a dedicated server which in the US seem to start around $80-100/month (just checking on serversearcher.com), DO and so on do provide a "dedicated cpu" product if that's too much.
- cess11 1y agoI've left a job because it was impossible to explain this to an ex-Googler on the board who just couldn't stop himself from trying to be a CTO and clownmaker at the company. The rough part was that we had made hardware investments and spent almost a year setting up the system for HA and immediate (i.e. 'low-hanging fruit') performance tuning and should have turned to architectural and more subtle improvements. This was a huge achievement for a very small team that had neither the use nor the wish to go full clown.
- exe34 1y agoI love that you're not just preaching - you're offering the service at a lower cost. (I'm not affiliated and don't claim anything about their ability/reliability).
- torginus 1y agoYup, I hope to god we are moving past the age of 'everything's fast if you have enough machines' and 'money is not real' era of software development. I remember the point in my career when I moved from a cranky old .NET company, where we handled millions of users from a single cabinent's worth of beefy servers, to a cloud based shop where we used every cloud buzzword tech under the sun (but mainly everything was containerized node microservices). I shudder thinking back to the eldritch horrors I saw on the cloud billing side, and the funny thing is, we were constantly fighting performance problems.
- bombcar 1y agoMy conspiracy theory is that "cloud scaling" was entirely driven by people who grew up watching sites get slash dotted and thought it was the absolute most important thing in the world that you can quickly scale up to infinity billion requests/second.
- colechristensen 1y agoNo, cloud adoption was driven by teams having to wait 2 years for capex for their hardware purchase and then getting a quarter of what they asked for. You couldn't get things, people hoarded servers they pretended to be using because when they did need something they couldn't get it. Management just wouldn't approve budgets so you were stuck using too little hardware. On the cloud it takes five seconds to get a new machine I can ssh into and I don't have to ask anyone for the budget. You can save a lot of money with scaling, you have to actually do that though and very few places do.
- dinvlad 1y agoAnd now, on cloud it’s the same but much more expensive and worse performance. We’ve been struggling for over a month to get a single (1) non-beefy non-GPU VM allocated on Azure, since they’ve been having insane capacity issues, to the extent that even “provisioned” capacity cannot be fulfilled ;-(
- 1y ago
- epistasis 1y agoWhat do you recommend for configuration management? I've had a fairly good experience with Ansible, but that was a long time ago... anything new in that pace?
- dijit 1y ago"new", I'm not sure, but I deployed 2,500 physical Windows machines with SaltStack and it worked pretty good. it also handled some databases and webservers on FreeBSD and Windows, I considered it better than Ansible.
- lazyfanatic42 1y agohaha this reminds me of when I used to manage Solaris system consisting of 2 servers. Sparc T7, 1 box in one state and 1 box in another. No load balancer. Thousands and thousands of users depending on that hardware. Extremely robust hardware.
- api 1y agoHow do you deprogram your devs and ops people from the learned helplessness of cloud native ideology? I've found that it's almost impossible to even hire people who aren't terrified of the idea of self-hosting. This is deeply bizarre for someone who installed Linux from floppy disks in 1994, but most modern devs have fully swallowed the idea that cloud handles things for them that mere mortals cannot handle. This, in turn, is a big reason why companies use cloud in spite of the insane markup: it's hard to staff for anything else. Cloud has utterly dominated the developer and IT mindset.
- awestroke 1y agoSo you'd rather self host a database as well? How do you prevent data loss? Do you run a whole database cluster in multiple physical locations with automatic failover? Who will spend time monitoring replication lag? Where do you store backups? Who is responsible for tuning performance settings?
- 7bit 1y agoI really don't understand this comment. The cloud doesn't protect you from data loss or provide any of the things you named.
- baby_souffle 1y agoYes it does? For a fraction of a dollar per hour, AWS will give me a URI that I can connect to. On the other end is a postgres instance that already has authentication, backups handled for me. It's also backed by a storage layer that is far more robust than anything I can get together in my rented cage with my corporate budget.
- theideaofcoffee 1y agoHosting a database is no different than self-hosting any other service. This viewpoint hath what cloud wrought, this atrophying of the most basic operational skills, as if running these magic services are only achievable by the hyperscalers who said they are the only ones capable. The answers to all of your questions are a hard: it depends. What are your engineering objectives? What are your business requirements? Uptime? Performance? Cost constraints and considerations? The cloud doesn't take away the need to answer these questions, it's just that self-hosting actually requires you to know what you are doing versus clicking a button and just hoping for the best.
- nikodunk 1y agoIf you’re big, invest in this. If you’re small, slap Dokploy/Coolify on it.
- rixed 1y agoI do not disagree, but just for the record, that's not what the article is about. They migrated to Hetzner cloud offering. If they had migrated to a bare metal solution they would certainly have enjoyed an even larger increase in perf and decrease in costs, but it makes sense that they opted for the cloud offering instead given where they started from.
- lazystar 1y ago> In reality, there is significant latency between Hetzner locations that make running multi-location workloads challenging, and potentially harmful to performance as we discovered through our post-deployment monitoring. the devil is in the details, as they say.
- rgrieselhuber 1y agoWe moved DemandSphere from AWS to Hetzner for many of the same reasons back in 2011 and never looked back. We can do things that competitors can’t because of it.
- dhruv_ahuja 1y agoCan you please explain what are some of those things? Curious to know and learn.
- dinvlad 1y agoThe cloud is also not fulfilling its end of the promise anymore - capacity on-demand. We’ve been struggling for over a month to get a single (1) non-beefy non-GPU instance on Azure, since they’ve been having just insane capacity issues, where even paying for “provisioned” capacity doesn’t make it available.
- traceroute66 1y ago> on S3 that is $20-$250 _per second_ on API calls! It is worth pointing out that if you look beyond the nickle & diming US-cloud providers, you will very quickly find many S3 providers who don't charge you for API calls and just the actual data-shifting. Ironically, I think one of them is Hetzner's very own S3 service. :) Other names IIRC include Upcloud and Exoscale ... but its not hard to find with the help of Mr Google, most results for "EU S3 provider" will likely be similar pricing model. P.S. Please play nicely and remove the spam from the end of your post.
- ksec 1y agoWe will soon have 256 Zen 6c per socket, so at least 512 Core per server. Multiple PCIe 5.0 SSD at 14GB/s at up to half a Petabytes storage, and TBs of Memory. And now Nvidia is in the game for Sever CPU, much faster time to market for PCIe in the future, and better x86 CPU implementation as well as ARM variants.
- themafia 1y ago> We typically see a doubling of performance The AWS documents clarify this. When you get 1 vCPU in a Lambda you're only going to get up to 50% of the cycles. It improves as you move up the RAM:CPU tree but it's never the case that you get 100% of the vCPU cycles.
- everfrustrated 1y agoSort of. 1 vcpu on x86 = 1 hyperthead not 1 core so yes you can't do uninterrupted work without some CPU "stolen". However that's not due to aws overhead or oversubscription but x86 architecture. For production workloads 2 vcpu should be minimum recommendation. On ARM where 1 vcpu = 1 core it is more straightforward.
- up2isomorphism 1y agoHalf year later all the data gets wiped out and what your customer can do? And you are still charging half of AWS, which is that case I am just doing these work myself if I really think AWS is too expensive.