4 ms·
I liked the clarity of the article and the reasoning is solid. Still, I'm not too convinced: - A key reason for splitting up a monolith into services is becaus
by atomicity 7y ago
I liked the clarity of the article and the reasoning is solid. Still, I'm not too convinced:
- A key reason for splitting up a monolith into services is because collaboration becomes too costly with 100s to 1000s of developers working on the same code. Your build & test system can't handle the number of commits. Your code takes too long to build. Would the data-access layer not them become the development bottleneck in such scenarios, meaning that it doesn't scale as well as SOA?
- What are the advantages and disadvantages of centralizing in the data-oriented layer over centralizing through a network-related tool like Envoy, Kubernetes, or Istio?
- The database layer itself often becomes a performance bottleneck, which requires us to run a sharded database. Some companies go the extra distance by running in-memory databases, document databases, and time-series databases. In such cases, wouldn't the data access layer need to support federation, which is a hard problem according to database research?
- Is O(N^2) really that big of a problem? It seems like the problem can be reduced to something simpler: developers cannot easily understand which services communicate with one another. If that is the case, would a visualization tool be sufficient?
- jeffffff 7y agoyes, at scale this turns the database into the bottleneck. this is not strictly a downside though, as it means you can centralize ownership of your database to one team of experts who can handle optimization, capacity planning, sharding, multi-tenancy, security, monitoring, etc for everyone. most product teams do not and should not need people with that expertise, so if you have multiple products or services running into scalability issues this approach can be a far more cost effective way of solving them than having each product or service handle these issues independently. while they do not use the term "data-oriented architecture", many of the largest web companies use what is effectively this approach and have teams dedicated to building and maintaining a shared data layer. some examples: google - spanner youtube - vitess, migrated to spanner facebook - tao uber - schemaless dropbox - edgestore twitter - manhattan linkedin - espresso notably absent is amazon. amazon has taken the full blown microservices approach where anyone can do whatever they want. worth noting is that amazon is in a very sad place when it comes to data warehousing and analyzing data across teams/products/etc. while the shared database approach is strictly intended for OLTP use cases and explicitly not meant for OLAP use cases, having a common interface to all data and something approaching a data model makes it extremely easy to replicate all your data out into a data warehouse or data lake or whatever you want to call your system for your OLAP workloads. with the 'every service has its own database' model, each team has to be responsible for replicating their data to analytics systems, and that is usually not super high on their priority list relative to product features. this problem is magnified when people from a different team want to consume data from that team's product/service but the team producing the data has no incentive to make it available. in large organizations (including amazon) this is a huge issue for teams who mostly do analysis, reporting, marketing, and other activities where they primarily consume data produced by others.
- pm90 7y agoChoosing an architectural design simply because it makes data warehousing easier doesn’t seem like a good enough reason to me. You give examples of all the Big Tech having such shared DBs but that seems like more of a reason to not use that pattern. Good DBAs are hard to find and not many people choose to become DBAs anymore. Big Tech can hire the experienced ones since they can compensate them pretty well; most companies can’t. The shared DB therefore becomes a critical bottleneck to the business.
- jeffffff 7y agobeyond some fairly large size of company it's less that it makes data warehousing easier and more that it makes centralized data warehousing possible. fortunately this type of environment is available today as a managed service in a few different offerings. gcp has spanner and vitess is available as a managed service on multiple cloud providers from planetscale.
- closeparen 7y agoCentralized data warehousing is possible as long as you constrain the number of distinct database engines and provide connectors for those. Services having private databases doesn't preclude data warehousing. It's why we have data warehousing! To enable joins across data from different silos.
- closeparen 7y agoThose things are database engines. Services can and do get their own instances. What those managed storage teams provide is akin to Amazon RDS, not one big database.
- jeffffff 7y agoyes they are database engines, but in many if not most of these cases there is only a single instance that is shared across all products at the company. it is very different than the rds model. of course there are access controls and abstractions such as schemas and tables but there aren't silos between data from different services
- Eyas 7y ago> Would the data-access layer not them become the development bottleneck in such scenarios, meaning that it doesn't scale as well as SOA? In terms of development, no, the data layer code grows sublinearly with the with the size/breadth of the schema/data. The data access layer is not much more than a database (plus usually, to enable event-driven programming, some semblance of subscriptions/notifications when data in your query changes). But it's fairly generalizable, and doesn't depend on the size of the team or schema using it. > Is O(N^2) really that big of a problem? It seems like the problem can be reduced to something simpler: developers cannot easily understand which services communicate with one another. If that is the case, would a visualization tool be sufficient? It really depends on how complex the system is. At some point, a visualization stops being helpful. There's obviously room to simplify a SOA dependency graph to look reasonable, and many do this successfully. But DOA is another interesting option in the toolkit: turn the problem on its head and say: maybe there's no graph at all.
- collyw 7y ago- A key reason for splitting up a monolith into services is because collaboration becomes too costly with 100s to 1000s of developers working on the same code. Your build & test system can't handle the number of commits. Your code takes too long to build. Would the data-access layer not them become the development bottleneck in such scenarios, meaning that it doesn't scale as well as SOA? You can't actually remove complexity like this, just push it to the dev ops layer. And also it makes setting up a development environment a lot more difficult.