4 ms·
Not central to the main ideas of this article, but if you want to have a data mesh that is self-service, why force folks to use a particular storage medium like
by bitsondatadev 5y ago
Not central to the main ideas of this article, but if you want to have a data mesh that is self-service, why force folks to use a particular storage medium like a data warehouse? That still requires centralization of the data.
Why not instead have a tool like Trino (https://trino.io https://trino.io) that allows you to let different domains use whatever datastore they happen to use. You still would need to enforce schema, but this can be done in tools like schema registry as mentioned in the article along with a data cataloging tool.
These tools facilitate the distributed nature of the problem nicely and encourage healthy standards to be discussed and the formalized in schema definitions and catalogs that remove the ambiguity of discourse and documentation.
Nice example is laid out in this repo of how Trino can accomplish data mesh principles 1 and 3 (https://github.com/findinpath/trino_data_mesh https://github.com/findinpath/trino_data_mesh).
- verbbis 5y agoFew data mesh proponents ”force” a particular storage medium - and the concept is largely agnostic regarding to this. But lots of early implementations in the wild have decided to standardize on it - either on some cloud object storage or, indeed, a cloud DW. One cannot argue how much it simplifies things in terms of manageability, access, cataloguing, performance… in an already complex architecture. Especially since no reference implementations exist. I understand that if your persistence layer is heterogeneous from the get go, layering on top of it might be a solution. But it is also an additional layer that needs to be managed. Conversely, in your opinion, what would be the shortcomings on centralizing on a modern, cloud-native data warehouse (tech, not the practice)? I see this being articulated less often.
- bitsondatadev 5y agoYou say, “one cannot argue performance of a data warehouse” but that’s precisely the issue with a DW. DW requires a lot of work to move data from the way that domains model their data on the service layer to how data is modeled in a central DW. You have to wait for data to become live to even begin running analysis on it. Setting up and worst of all maintaining pipelines is an expensive undertaking in both time and money. It’s not to say the DW is bad and never the solution. The problem is making it the only solution and not providing domains the flexibility to model data the way they need it. You say it’s more complex to manage but that’s the idea behind data mesh, you don’t manage that part, the team with their domain knowledge and data solution does. They can make it as simple or complex as they want internally but if they follow the standards to play in your data mesh who cares? Not your problem. For example say a domain needs realtime data analytics and use something like Druid to store their data. That’s fine. If they want to play in the data mesh you’ve provided, they just need to follow the rules in their data model, but they don’t need to use a cloud DW to do that. You can’t argue that avoiding the copying of terabytes of data a day from a domain to a DW is more performant than adhoc analysis (MB to GB) of that data. Why move or copy a dataset when you don’t need to? Why force domains to use any solution that’s not actually solving their domain problem?
- verbbis 5y agoI have to disagree with almost all of the points you raise. The degree to which you wish to (re)model data is a design decision - also within a DW. This is what I mean by divorcing the tech from the practice. Methods to store and manipulate semi-structured data exist within cloud data warehouses and there is nothing inherent in the technology preventing from extending said support. Also, it seems to me, that even in advanced analytics incl. ML one eventually works with the data in a tabular form anyway. When it comes to performance, I was referring primarily to e.g. cross-domain queries. That continues to be challenging in data virtualization / federated query engines. You (and data mesh) focus on the domains - and nothing prevents carving out ample portions of storage and compute from a cloud DW to them and doing the exact things you propose. Except that "federated computational governance" and enforcing standards is now heaps easier since everyone is relying on the same substrate. Data mesh targets analytical workloads - and surely you do not suggest e.g. hooking Trino up directly to operational, OLTP databases? One has to, as least how things stand today, copy the data somewhere anyway for not the least historical analysis, as well as often transform it to be understandable to a downstream (data product) consumer. So you will have "pipelines" no matter what - even in a data mesh. Not only are the domain teams free to build the aforementioned pipelines and model data in any way they see fit - within a cloud-native DW, but also the DataOps tooling available in this area is already relatively mature. The only part I can somewhat relate to is about "forcing" to rely on a specific piece of technology. But I think that is just something one needs to accept in a corporate setting and a balancing act. And I'm talking about the majority of the regular companies out there, not FAANG. Also, there other arguments to be made - such as avoiding vendor lock-in or if the capabilities of the DW simply does not cater to the specific problem your domain has. But these are not the arguments you made.
- bitsondatadev 5y agoAgree to disagree then I guess. "When it comes to performance, I was referring primarily to e.g. cross-domain queries. That continues to be challenging in data virtualization / federated query engines." Have you tried Trino lately? Data virtualization like Denodo still relies on moving data back between engines to execute a query, Trino pushes the queries down to both systems and processes the rest in flight. The fact that you use both of those interchangeably makes me think you may not have tried it. "Data mesh targets analytical workloads - and surely you do not suggest e.g. hooking Trino up directly to operational, OLTP databases?" Not operational data. Are you saying teams aren't allowed to store immutable data in PostgreSQL? "Not only are the domain teams free to build the aforementioned pipelines and model data in any way they see fit" You will organically see teams build data infrastructure with different tooling. This happens when you don't tell a team they have to use a DW to play the analytics game. I have never seen a company that has multiple teams that organically land on using the same tech. So naturally (that is without forcing the to use one substrate i.e. DW) you will have different teams (or domains) using different databases. This is why we have DW to begin with. They were literally created to be a copy of domain data that naturally lied in many operational databases. I get that in an ideal world, there would be some magical one size fits all solution that everyone would just use. However, that system doesn't exist. It's certainly not the DW. DW can be one of those solutions, just not THE one.