8 ms·
Working in realtime games that require high bandwidth usage, AWS and basically every other public cloud is fundamentally unusable for us because of this exact p
by AgentK20 5y ago
Working in realtime games that require high bandwidth usage, AWS and basically every other public cloud is fundamentally unusable for us because of this exact problem. Our bare metal infrastructure for the Hypixel Minecraft network uses 3-4PB per month, so we simply lease a 100gbps transit link billed at 95th percentile.
Last I checked, AWS wanted ~$200k per MONTH, for just bandwidth (no compute, memory, storage, or anything else). We'd love to be able to use the cloud, but not at the cost of increasing our monthly expenses by an entire order of magnitude, so we just stick to bare metal colocation.
I honestly believe that if you have the technical skill in-house, and are spending more than, say, $30k/mo on public cloud hosting, that you should seriously evaluate whether bare metal could significantly decrease your costs.
- aclelland 5y agoAt $JOB, we have a similar issue with static content. We serve over 1PB a month and the price AWS would charge to use CloudFront is an order of magnitude larger than even a Cloudflare Enterprise plan (which comes with some nice bells and whistles that Cloudfront doesn't offer). Even with discounts from AWS it just doesn't make sense to use AWS to serve up the assets to users. We do use S3 as our static asset backend and the combination works really well. I would love to see Cloudflare release a S3 compatible storage service though. I think we'd jump onto that in a heartbeat.
- HeavenFox 5y agoBackblaze has S3-compatible API and they have free data transfer to Cloudflare
- aclelland 5y agoThe last time I looked at BB S3 API they didn't offer the lifecycle controls that AWS over. Mainly the ability to remove old versions of files after a period of time and we didn't want to roll our own expiry solution. Might be worth looking again though.
- treesknees 5y agoYep, B2 does offer this capability now. They've been working hard to add S3-compatible/similar features. "Lifecycle rules instruct the B2 service to automatically hide and/or delete old files. You can set up rules to do things like delete old versions of files 30 days after a newer version was uploaded." https://www.backblaze.com/b2/docs/lifecycle_rules.html https://www.backblaze.com/b2/docs/lifecycle_rules.html
- artbackblaze 5y agoArt from Backblaze here...Thanks for the feedback! As treesknees points out, we currently support lifecycle rules with our B2 Native API and our web application - those are available right now if you need to do it ASAP. https://www.backblaze.com/b2/docs/lifecycle_rules.html https://www.backblaze.com/b2/docs/lifecycle_rules.html
- 0xy 5y agoCloudFront is also a pretty inferior product in the CDN marketplace. Why do you think Amazon's retail business signed a huge contract with Fastly to be their primary CDN for all mission critical retail images? CloudFront sucks. Not competitive on price, features, performance or anything other than "we already pay AWS so might as well use it".
- Rd6n6 5y agoWhat is the best CDN offering right now in your opinion?
- 0xy 5y agoIf you want raw speed/features/reach, then Fastly or Akamai. CloudFlare's CDN is competitive on features and price but it isn't a market leader in terms of raw speed.
- manigandham 5y agoRaw speed in what way?
- manigandham 5y agoCloudflare. Their performance, pricing, and Workers platform means you can build whatever you need.
- acj 5y agoOne advantage of CloudFront is support for uploading large files (5GB+) to the origin server. CDNs tend to enforce a low size limit for uploads. It’s probably an uncommon use case, but the reduced client-side latency is nice for customers who are far away from the origin.
- dvaun 5y agoWhat if your business has fluctuating loads? I can see how running game servers—with a (somewhat, maybe?) predictable load of players—can be done efficiently on baremetal and with colocating. For another business that has huge peaks of demand, such as analytics with dynamic queries, I fail to see how baremetal can compare to spinning up hundreds of instances on-demand. Perhaps it comes down to what services you offer your clients and how you implement them?
- benlivengood 5y agoJust compare the spot/preemptible instance price or committed-use price to the on-demand price to see how it can be cheaper.
- brianwawok 5y agoThe scale part of the cloud is often not as exciting as it seems on paper. Yes a cloud can auto-scale 100x. But can everything else support that? Like is your DB setup to handle 100x increase in load? (Sure, use dynamo DB, but that has other restrictions). If the cloud to bare metal price difference is 10x.. you could easily just buy bare metal = 2x your peak load, and still come out ahead..
- dvaun 5y agoWhat you state makes sense. For me, I haven't worked in an environment (yet) that would need to handle fluctuating loads at scale—so my comment is my own speculation based on my experience working for smaller businesses with MUCH less data and bandwidth usage compared to those mentioned here. And that's why I come to HN, lobsters, etc :) so that I can read and learn from others' experiences...
- freedomben 5y agoI've worked with dozens of different companies, and it's pretty rare to truly have such crazy variance in load.
- 5y ago
- deleted 5y ago[deleted]
- cblconfederate 5y ago> We'd love to be able to use the cloud why?
- gyoza 5y agobecause you are not limited by your fleet.
- walrus01 5y agoFor pricing reference you can get a 100GbE transit link (from a top-20 sized carrier by CAIDA ASrank size) at major IX points now for well under $6000 a month. And if you are present with your own bare metal infrastructure at such a place you almost certainly also have the opportunity to connect to a serious IX for settlement-free open peering, and to run PNIs to other major sources or sinks of your traffic. So by no means will all of your traffic be going through transit.
- throw_nbvc1234 5y agoHow much would 1k 100GbE links cost from one of those providers? How much would 10k of those links cost? Does the price increase linearly or exponentially?
- walrus01 5y agoIf you need multiple 100GbE transit connections from ISPs larger than yourself, in multiple locations, you most likely also are a fair sized ISP, so the situation is very different because you'll also be purchasing a variety of transport (point to point circuits) between cities at 100/200/400GbE capacity, various DWDM circuits, lighting your own dark fiber, etc. It's a whole other ball game.
- bombcar 5y agoIt will be linear until you reach a percentage of the capacity of the provider themselves. Not every company has multiple terabits into each datacenter. At that point you’re probably buying dark fiber to get where you need to go.
- Danieru 5y agoThis is super true, yet sadly much of the Japanese mobile gaming industry is paying through the nose for cloud hosting. There just is not a culture of optimizing the cost structure.
- Zababa 5y agoIsn't this because they're basically printing money with slightly disguised gambling?
- ketzo 5y agoAre they printing money with gacha games? Yes. Doesn't mean they couldn't do some optimization! But to your point, yeah, I imagine there's probably a little more pressure to develop new features/characters/etc. rather than spend months rigging up in-house network infrastructure.
- late2part 5y agoAre you talking about Amazon?
- Zababa 5y agoNo, I'm talking about Japanese mobile game companies and their gacha games.
- Zababa 5y agoI think a big driver for adoption of AWS and the cloud in general is tech salaries in the US. You'll often hear people on HN talking about how useless it is to try to save $Xk because the engineer cost will be even higher than that.
- deleted 5y ago[deleted]
- kelp 5y agoThis is true until you get to a certain scale and then you start looking at your margins and realize a ton of it is your wasteful use of AWS. I spent a big chunk of my last 2 jobs dealing with that issue.
- Zababa 5y agoThat's true, and on the opposite side you have companies that manage so much hardware that it starts being profitable for them to become a cloud operator.
- kazen44 5y agoalso, the math for some cloud providers doesn't work that well if you are in regions which are not EU-west or the US. even in europe, AWS has "meh" regions compared to azure or many of the colo/dc locations.
- fwsgonzo 5y agoIs there any cacheable content? I'm working on high-performance compute on CDN software you can run yourself.
- AgentK20 5y agoNot when we don't control the game client. We're the largest Minecraft server in the world (220k peak concurrent users) yet we operate entirely as a third party with no involvement from Microsoft.
- JMTQp8lwXL 5y agoThere's the opportunity cost to throwing away your organization's knowledge of AWS, though. Everyone will have to learn to do things the bespoke way on your bare metal. So developer productivity could decline.
- freedomben 5y agoThis is where Kubernetes I think helps a ton. I work with bare metal customers all the time that stand up OpenShift on their BM and can migrate k8s apps very easily. Depending on the apps you may need to throw an object storage solution in there too (such as OpenShift Container Storage). It does require a certain scale before this makes sense, but it's not nearly as high of a scale as most people think.
- surfer7837 5y agoYou go on to AWS for their managed services like Fargate, DynamoDB, ECS, S3 etc. Have used OpenShift in the past and had endless problems with cluster stability (especially in 3.x), and weird inconsistencies. With AWS I could just spin up 10 Kubernetes clusters with pretty much unlimited resources, can't do that in OpenShift because you'd hit a resource quota or limit.
- AgentK20 5y agoI would say the exact opposite in some ways for teams who are not using AWS yet. My team for example has been operating bare metal for 8 years now, and we know how to do that. Transitioning our team to the "bespoke way" that AWS does things has a huge opportunity cost, too.
- late2part 5y agoThis is the correct answer. People are swimming in the koolaid and don't think a different environment exist. These kids. When I was a kid......
- dilyevsky 5y agoImo using cloud vendor apis is more bespoke than using kube/nomad or even old school andible/salt on baremetal cluster. With the exception of s3 all your knowledge will tabled if ever need to switch the vendors
- ianhawes 5y agoIs *that* why Hytale hasn't been released yet?
- jordo 5y agoThis is exactly why we are using DigitalOcean to host our gameservers... Our egress costs are basically 1/10th of what we would pay on the big three. DigitalOcean Egress - $0.01/GB
- agucova 5y agoHow do you orchestrate or manage your servers currently? And do you employ a in-house solution for it as well?
- zxcvbn4038 5y agoI’ve found that are generally three stages in a company’s life. If you are extremely technical you can make bare metal beat cloud in the beginning. As the environment grows and IT roles diversify cloud starts to make more sense financially then bare metal so you do a cloud migration. If you are lucky then eventually you get to the point that it’s cheaper to run your own hardware then rent someone else’s and you end up migrating out of the cloud again. If there is a fourth stage then it’s getting to the point that it’s more profitable to run a side business renting out your unused resources to others and you become a cloud. I’m surprised there isn’t a Walmart cloud or an Exxon cloud by now.
- thaumasiotes 5y ago> If there is a fourth stage then it’s getting to the point that it’s more profitable to run a side business renting out your unused resources to others and you become a cloud. This is a common belief about AWS, but I find the counterpoint persuasive - if AWS involved Amazon renting out its unused resources, it would have to shut down in December. But obviously AWS does not shut down in December. They're not renting out their unused resources; they're renting out their experience running this kind of system.
- px1999 5y agoThe fourth stage doesn't even involve owning your own infrastructure. We pay a 3rd party for our AWS (and Azure) usage to get better rates than AWS would give us (and without a commitment to increased spend like AWS reps have tried to push for Amazon's available programs). I wouldn't be surprised if they just used volume discounts / negotiated pricing and reservations at their end to make it massively profitable for them.
- zxcvbn4038 5y agoI think SPOT (pre-emptable) instances are an under appreciated resource in AWS. Technically they can get pulled back at any time but in practice it’s only happened to me once, for a single instance. I’ve lost far more instances to hardware failures in the same period. If you can get your head around that then you can save more money with SPOT plans then any of the regular compute plans. AWS also does an auction for SPOT instances which can push the price really low for ARM and previous generation instances whereas other clouds just give you a set discount when using them.
- pugz 5y agoI think it depends very much on what your product does. E.g. in your case (games) bandwidth is a big deal. The last three companies I've worked at have all had AWS bills of more than $10M/yr and data egress costs were never more than 2-8% of the bill. A mixture of B2C and B2B SaaS companies. In that case, the economics of bare metal aren't as clear-cut.
- late2part 5y agoThis is such a great conversation, and I agree with you - that while the egress is egregious; it's rarely a large portion of the cost. I run a team that manages our public and private (DC) cloud for a successful tech company. - I need more folks to help us do things better. Please contact me Mr. Pugs or others if you are open to thinking about looking at another company.
- MrStonedOne 5y agoI run servers for one of the biggest open source multiplayer games on github. (tgstation/tgstation) Everybody asks why we don't do cloud, and time and time again i have to point out that the fastest per-thread cpu on ec2 costs more per month then our rented bare metal server, that benchmarks as faster per thread, and has 3x the bandwidth we need, included in the price!. Even on a cpu level, the obsession with intel """enterprise""" cpus is leaving money on the table. "but what about web stuff, surely that could go on the cloud!" the best cpus to run (mostly) single threaded game servers right now come with 8 to 16 cores, and we aren't running that many game servers in each region we target, so what else are we gonna do with those spare cores but load up ansible managed docker swarms to handle the webstuff? Notice how I haven't even needed to get to the egress cost to point out how senseless going to the cloud can be.
- stewartmcgown 5y agoI get very happy seeing people run Docker Swarm in production. we need a community!
- TheFlyingFish 5y agoI run Docker Swarm on my homelab, but I honestly wouldn't touch it for production because it's lacking features that I find critical and doesn't show any signs of developing them soon. For example, Swarm secrets can only be exposed as files, not as environment variables. You can argue this is more secure because file permissions are more granular than env vars, but IMO that's a silly argument in a container context because containers are almost always single-user to begin with. Moreover, the vast majority of containerized applications expect their secrets in env vars, so you have to resort to fragile entry point wrappers if you want to make use of Swarm secrets.
- MrStonedOne 5y agohonestly the biggest failure of swarm secrets is the inability to change them while things are using them. automating secret usage across containers that are deployed via ansible scripts gets annoying in that situation.
- late2part 5y agoYou, sir, are smart, don't believe hype, and know how to math. You are not AWS's target market.
- ckdarby 5y agoIn the past I worked at a company in the top 20 bandwidth consumer in the world. 95th billing is the way to go. Due to the nature of the company's content (adult) nobody wanted to do peering agreements due to the public outcry if anyone found out but if that had been possible I figure the plan would have been to follow Netflix's open connect journey. How to run a top 20th largest bandwidth consumer in the world for < $1M/month. https://openconnect.netflix.com/en/ https://openconnect.netflix.com/en/ Disclaimer: These are only my own personal opinions