4 ms·
yes, at scale this turns the database into the bottleneck. this is not strictly a downside though, as it means you can centralize ownership of your database to
by jeffffff 7y ago
yes, at scale this turns the database into the bottleneck. this is not strictly a downside though, as it means you can centralize ownership of your database to one team of experts who can handle optimization, capacity planning, sharding, multi-tenancy, security, monitoring, etc for everyone. most product teams do not and should not need people with that expertise, so if you have multiple products or services running into scalability issues this approach can be a far more cost effective way of solving them than having each product or service handle these issues independently.
while they do not use the term "data-oriented architecture", many of the largest web companies use what is effectively this approach and have teams dedicated to building and maintaining a shared data layer. some examples:
google - spanner
youtube - vitess, migrated to spanner
facebook - tao
uber - schemaless
dropbox - edgestore
twitter - manhattan
linkedin - espresso
notably absent is amazon. amazon has taken the full blown microservices approach where anyone can do whatever they want. worth noting is that amazon is in a very sad place when it comes to data warehousing and analyzing data across teams/products/etc. while the shared database approach is strictly intended for OLTP use cases and explicitly not meant for OLAP use cases, having a common interface to all data and something approaching a data model makes it extremely easy to replicate all your data out into a data warehouse or data lake or whatever you want to call your system for your OLAP workloads. with the 'every service has its own database' model, each team has to be responsible for replicating their data to analytics systems, and that is usually not super high on their priority list relative to product features. this problem is magnified when people from a different team want to consume data from that team's product/service but the team producing the data has no incentive to make it available. in large organizations (including amazon) this is a huge issue for teams who mostly do analysis, reporting, marketing, and other activities where they primarily consume data produced by others.
- pm90 7y agoChoosing an architectural design simply because it makes data warehousing easier doesn’t seem like a good enough reason to me. You give examples of all the Big Tech having such shared DBs but that seems like more of a reason to not use that pattern. Good DBAs are hard to find and not many people choose to become DBAs anymore. Big Tech can hire the experienced ones since they can compensate them pretty well; most companies can’t. The shared DB therefore becomes a critical bottleneck to the business.
- jeffffff 7y agobeyond some fairly large size of company it's less that it makes data warehousing easier and more that it makes centralized data warehousing possible. fortunately this type of environment is available today as a managed service in a few different offerings. gcp has spanner and vitess is available as a managed service on multiple cloud providers from planetscale.
- closeparen 7y agoCentralized data warehousing is possible as long as you constrain the number of distinct database engines and provide connectors for those. Services having private databases doesn't preclude data warehousing. It's why we have data warehousing! To enable joins across data from different silos.
- closeparen 7y agoThose things are database engines. Services can and do get their own instances. What those managed storage teams provide is akin to Amazon RDS, not one big database.
- jeffffff 7y agoyes they are database engines, but in many if not most of these cases there is only a single instance that is shared across all products at the company. it is very different than the rds model. of course there are access controls and abstractions such as schemas and tables but there aren't silos between data from different services
- closeparen 7y agoI guess neither of us want to out our employment history here, but for the ones I know about, that's absolutely not true. Reading from another service's database is sometimes possible but always considered hacky tech debt.
- jeffffff 7y agofrom what i've heard from reliable sources there are only 2 spanner clusters, one for ads and one for everything else. i'd be surprised if there isn't an isolated one for gcp but for internal google products there are only 2. i also have it on good authority that data sharing between services through edgestore and tao is common. i have less insight into the others.