5 ms·
Disclaimer: I work on Google Cloud so I will be speaking from the bias of knowing those products. They talk a lot about reducing operating complexity and scali
by tweenagedream 9y ago
Disclaimer: I work on Google Cloud so I will be speaking from the bias of knowing those products.
They talk a lot about reducing operating complexity and scaling their infrastructure, I wonder what the cost of their current infrastructure + the staff to maintain it might be vs the managed solutions that cloud providers offer now.
For example, using cloud datastore or spanner or big table as a persistent layer, these managed services can definitely scale to the current need and I've seen them go much higher as well.
For logs ingestion and analysis, big query can be a very powerful tool as well, and with streaming inserts that data can be queried in near real time. For things that are less urgent, batch queries. For other things dataflow can help with streaming workloads.
I think one of the problems they alluded to though was that at the moment they're on a single provider, and what they're looking for is a multi cloud strategy which totally makes sense. A lot of the above products create some kind of locking, with some exceptions, like using hbase as an interface to big table or beam as an interface to dataflow. Though I don't know what the other providers offer that may have these same interfaces.
Another option is kubernetes, which I believe all providers are pretty strongly embracing. Having most of the supporting infrastructure be brought up with a few kubectl commands could help them scale across several cloud providers quickly.
- matt_s 9y agoI think they detailed in the article that the problem isn't their game servers which are AWS cloud based and can scale up, it is their login/setup/matchmaking (my term) server infrastructure that is the first thing users first encounter that is having issues. Usually there is a cost/scale threshold with managed providers where it is cheaper do DIY than to pay thousands upon thousands per month for say log ingestion.
- tweenagedream 9y agoI agree and my comment was mostly aimed at the supporting services rather than the actual game servers. However, they mentioned that they flew out Mongo experts to help on site and during weekend peaks. That sounds pretty expensive to me. Many times, the cost of the infrastructure is the cheap part, and the salaries of the humans managing and tuning it are the expensive part. I wonder if they really hit the point where doing this part in house rather than offloading it is a net win.
- ben_jones 9y agoCongratulations you are now a moderator of /r/mountainview
- trevyn 9y agoAgreed, Cloud Spanner is an impressive piece of technology, but question: If you build your business on Spanner and it does start having problems for whatever reason, what do you do? Obviously at this size you’d get great support from Google, but ultimately you rely on one managed provider and your hands are tied. That’s a tough situation to be in when you’re servicing 3.4M users.
- ShakataGaNai 9y agoAll of the managed products from the different cloud providers are, more or less, great. The problem is they are black boxes. When something goes wrong you're completely at the whim of support. Ever call Dell/Comcast support and want to tear your hair out? Yea... it's like that except neither AWS nor Google have phone numbers to call. The other problem is that most of these things aren't easy to migrate to. AWS RDS is much easier because its just managed whatever you're already using. But cloudspanner? DynamoDB? You have to completely re-architect your application. Then you have to move your application, and data, to this new system...without massive outages. It's a lot of work and a lot of cost. So until things go HORRIBLY sideways, most companies don't have the spare time/money. Been there, tried that.
- trevyn 9y agoIf your service is mission-critical, you’ll have a support contract, and there absolutely is a number you can call. I have high confidence that Google and Amazon will make every effort to make sure their services perform to spec, the real problem is the feeling of helplessness when something does go wrong.
- ShakataGaNai 9y agoOf course you have a support contract, but how fast you can get a response still doesn't change. AWS Business Support - which is what is very common, a "Production Down" has a 1hr SLA for response. That's a long ass time for a database to be down without doing anything more than sitting on your hands.
- nogbit 9y agoEpic would have enterprise with the money they are making. That's business critical <15min response time for production down. https://aws.amazon.com/premiumsupport/compare-plans/ https://aws.amazon.com/premiumsupport/compare-plans/
- bluesroo 9y ago> Yea... it's like that except neither AWS nor Google have phone numbers to call. At my current company we have access to AWS support. I'm not high enough on the ladder to know the specifics (separate contract, size threshold?), but when we've had issues in Aurora we've had personalized support. I have no doubt if you're scaling to thousands of instances you would have personalized support, if for no other reason that a cloud provider would not want to lose such a huge customer.
- tlynchpin 9y ago> .. currently unclear to us and support why our writes are being queued .. You think GCP offers better support on spanner et al when customer is having performance problems? In this case probably yes, because an Epic sized monthly spend is highly effective at escalating through support. It takes low effort to find <cloud persistence horror story> around here so we know cloud is not a special magic that is immune from integration performance problems. But the economic incentives are meaningfully different and especially so at runtime.