6 ms·
The question is how often is that necessary? Once again the point goes back to the article title. You are not Google. Unless your product is actually large, you
by jacoblambda 8y ago
The question is how often is that necessary? Once again the point goes back to the article title. You are not Google. Unless your product is actually large, you probably don't need all of that and even if you do, you can probably just do part of it in the cloud for significantly cheaper and get close to the same result.
This obsession with making something completely bulletproof and scalable is the exact problem they are discussing. You probably don't need it in most cases but just want it. I am guilty of this as well and it is very difficult to avoid doing.
- scarface74 8y agoYou think only Google needs to protect against data loss? We have a process that reads a lot of data from the database on a periodic basis and sends it to ElasticSearch. We would either have to spend more and overprovision it to handle peak load or we can just turn on autoscaling for read replicas. Since the read replicas use the same storage as the reader/writer it’s much faster. Yes we need “bulletproof” and scalability or our clients we have six and seven figure contracts with won’t be happy and will be up in arms.
- aprdm 8y ago> You think only Google needs to protect against data loss? Just have a cron job taking backups in a server and sending somewhere else? It has been working for the last 40 years...
- scarface74 8y agoSo a cron job can give me point in time recovery? Can a cron job give me automatic failover to another server that is in sync with the master? Can it give me autoscaling read replicas? Who is going to get up in the middle of the night when the cron job fails? Yes I am sure my company that has six and seven figure contracts would be just as well served self hosting MySQL on Linode.
- candiodari 8y agoOTOH the last "cloud" solution I've seen was a database that allowed a team to schedule itself. It was backed on Google cloud, had autoscaling backend. It was "serverless" (app engine), and had backups etc configured. Cost to run it for 4 years: $70000. QPS ? 5 was the highest I ever found. You don't need this capacity. You just don't. But it's got to be a good business to be in ...
- scarface74 8y agoInstead of just assuming pricing based on Google pricing (which we weren’t talking about), you could always just find AWS pricing. https://aws.amazon.com/rds/aurora/pricing/ https://aws.amazon.com/rds/aurora/pricing/ Storage cost of 500Gb of data for four years - $2400 Backing up is free up to the amount of online storage. IO request are $0.12 per million. Transfer to/from VMs in the same availability zone is free. A r5.large reserved instance is $6600 a year. Of course if your needs are less, you could get much cheaper cpu and memory.
- candiodari 8y agoIssue is, what would the charge be for a typical application for some basic business function, say scheduling attendance. So QPS really low, <5. But very spread out, as it's used during the workday by both team leader and team members. Every query results in a database query or maybe even 2 or 3. So every, say 20 minutes or so there's 4-5 database queries. Backend database is a few gigs, growing slowly. Let's say a schedule is ~10kb of data, and that has to be transferred on almost all queries, because that's what people are working on. That makes ~800 gig transfer per year. This would be equivalent to something you could easily store on say a linode or digital ocean 2 machine system, for $20 month or $240/year + having backup on your local machine. This would have the advantage that you can have 10 such apps on that same hardware with no extra cost. And if you really want to cheap out, you could easily host this with a PHP hoster for $10/year. So how do you calculate the AWS costs here ?
- scarface74 8y ago
- tigershark 8y agoWhy do you have a specific need to read the database periodically polling it instead of just pushing the data to elastic search at the same time that it reaches the database? I don’t know anything about your architecture, but unless you’re handling data on a very big scale probably rationalising the architecture would give you much more performance and maintainability than putting everything in the cloud.
- scarface74 8y agoWithout going into specifics. We have large external data feeds that are used and correlated with other customer specific data (multitenant business customers) and it also needs to be searchable. There are times we get new data relevant to customers that cause us to reindex. We are just focusing on databases here. There is more to infrastructure than just databases. Should we also maintain our own load balancers, queueing/messaging systems, CDN, object store, OLAP database, CI/CD servers, patch management system, alerting monitoring system, web application firewall, ADFS servers, OATH servers, key/value store, key management server, ElasticSearch cluster etc? Except for the OLAP database. All of this is set up in some form in multiple isolated environments with different accounts in one Organizational Account that manages all of the other sub accounts. What about our infrastructure overseas so our off shore developers don’t have the latency of connecting back to the US? For some projects we even use lambda where we don’t maintain any web servers and get scalability from 0 to $a_lot - and no there is no lock-in boogeyman there either. I can deploy the same NodeJS/Express, C#/WebAPI, Python/Django code to both lambda and a regular old VM just by changing my deployment pipeline.