5 ms·
They were probably looking for something non-relational, not suggesting you write your own.
by bowmessage 7y ago
They were probably looking for something non-relational, not suggesting you write your own.
- dijit 7y agoThe exact message I got back was "Attempted to use magic software solution", so I believe they intended for me to say "some kind of relational" or "some kind of non-relational" database and maybe some key criterion about what kind of access pattern instead of me internalising the problem and pulling out a ready-made solution.
- username90 7y agoWhen I worked on the infrastructure parts of Google I worked on systems with millions of QPS, does PostgreSQL really handle that kind of load gracefully?
- dijit 7y agoPostgres was just one part; I was describing a sharding solution that was using Postgres as long term storage underneath with a memory-based distributed message queue for ingestion and sharded cache layer for egress. Sure, PostgreSQL scales relatively nicely on single nodes but I chose it because it has a write-ahead log, strong transaction isolation and b-tree indexing, which would have been useful given the question I had.
- username90 7y agoDid you put all the data in a single postgres? Since a million queries per second is more than postgres can handle. And if you don't put them all in one database then what is the point of postgres features like transactions or indexing? At least on the teams I saw people did these things in code, or they used a solution some other engineer at Google had written, I don't think there are any public databases which handles these things nicely without costing ridiculous amounts to run. So yes, writing a database is a part of the responsibilities a SWE at Google could be expected to handle. And this isn't some kind of special role, the people working on these things were normal L3-L5 engineers. Edit: Note that there were SRE's working on these things as well, infrastructure teams are often mixes of SWE's and SRE's and their roles overlap somewhat, sometimes SRE's builds entire things themselves because they understand the production environment better.
- dijit 7y agoNo I did not put all my data into a single postgres, even if the TX/s would have scaled (they wouldn't have) the data volume would have exceeded the limits of what a single server can provide. My solution was dependent on splitting the data into sub-categories; for the bulk of the data I was going to use idempotent sharding based on a unique key, I said I would have implemented it as a SHA1 of a userID modulus'd by 512, with 512 being the upper bound on the number of shards/machines, (or a multiple of that; at the scale I was given it would have worked). I then went into detail about how much a single machine would need to ingest and my own experience with postgresql performance, I also spoke at length about what the maximum theoretical volume of data was for a single DC (however, that was "not important" the recruiter indicated I had a magic datacenter that did not have problems with cross-connecting many, many hundreds of GB/s in a mesh). Frankly, I already build global solutions in my day job, sure they're not google scale, but they're built to order, quite cost effective and what's more important: they function very well and are engineered to the point where we know beyond reasonable doubt that they will perform as needed on day 1. (I work with always-online video games, the first day is the worst day, scalability wise)
- username90 7y ago> the recruiter indicated I had a magic datacenter that did not have problems with cross-connecting many, many hundreds of GB/s in a mesh Well, then this is different than your original description, I'd need to get more details about the problem but he is right that machine to machine connections in a data center doesn't scale very well. This might not be a problem at the scales you are used to but it is a problem at Google scale. This is a very common problem that is not obvious at first when you work with data centers, I guess he just assumed that you would know this. Knowing your background you would probably adapt to it quickly on the job, but I guess they just asked the same question to every experienced SRE they got? Edit: Another problem with your solution is that you used a static sharding strategy and didn't consider that increasing demand in the future would force you to reshard the database. Downtime might be accepted in the video game industry, and there you most likely wont even get much more demand than day 1, but using sharding strategies which lets you reshard in real time without downtime is more or less a must on the projects I worked on.