6 ms·
Great talk, there's a lot I can relate to in here. I find this topic difficult to navigate because of the many trade-offs. One aspect that wasn't mentioned is
by loevborg 1y ago
Great talk, there's a lot I can relate to in here.
I find this topic difficult to navigate because of the many trade-offs. One aspect that wasn't mentioned is temporal. A lot of the time, it makes sense to start with a "database-oriented design" (in the pejorative sense), where your types are just whatever shape your data has in Postgres.
However, as time goes on and your understanding of the domain grows, you start to realize the limitations of that approach. At that point, it probably makes sense to introduce a separate domain model and use explicit mapping. But finding that point in time where you want to switch is not trivial.
Should you start with a domain model from the get-go? Maybe, but it's risky because you may end up with domain objects that don't actually do a better job of representing the domain than whatever you have in your SQL tables. It also feels awkward (and is hard to justify in a team) to map back and forth between domain model, sql SELECT row and JSON response body if they're pretty much the same, at least initially.
So it might very well be that, rather than starting with a domain model, the best approach is to refactor your way into it once you have a better feel for the domain. Err on the side of little or no abstraction, but don't hesitate to introduce abstraction when you feel the pain from too much "concretion". Again, it takes judgment so it's hard to teach (which the talk does an admirable job in pointing out).
- bubblyworld 1y agoPretty naive question, but what differentiates a "domain model" from these more primitive data representations? I see the term thrown around a lot but I've never been able to grok what people actually mean. By domain model do you mean something like what a scientist would call a theory? A description of your domain in terms of some fundamental concepts, how they relate to each other, their behaviour, etc? Something like a specification? Which could of course have many possible concrete implementations (and many possible ways to represent it with data). Where I get confused with this is I'm not sure what it means to map data to and from your domain model (it's an actual code entity?), so I'm probably thinking about this wrong.
- metayrnc 1y agoFor me domain model means capturing as much information about the domain you are modeling in the types and data structures you use. Most of the time that ends up meaning use Unions to make illegal states unrepresentable. For example, I have not seen a database native approach to saving union types to databases. In that case using another domain layer becomes mandatory. For context: https://fsharpforfunandprofit.com/posts/designing-with-types-making-illegal-states-unrepresentable/ https://fsharpforfunandprofit.com/posts/designing-with-types...
- skydhash 1y agoA quick example can be found with date. You can store it in ISO 8601 string and often it makes more sense as this is a shared spec between systems. But when it comes to actually display it, there's a lot of additional concerns that creep in such as localization and timezones. Then you need to have a data structure that split the components, and some components may be used as keys or parameters for some logic that outputs the final representation, also as a string. So both the storage and presentation layer are strings, but they differs. So to reconcile both, you need an intermediate layer, which will contains structures that are the domain models, and logic that manipulate them. To jump from one layer to another you map the data, in this example, string to structs then to string. With MVC and CRUD apps, the layers often have similar models (or the same, especially with dynamic languages) so you don't bother with mapping. But when the use cases becomes more complex, they alter the domain layer and the models within. So then you need to add mapping code. Your storage layers may have many tables (if using sql), but then it's a single struct at the domain layer, which then becomes many models at the presentation layer with duplicate information. NOTE That's why a lot of people don't like most ORM libraries. They're great when the models are similar, but when they start to diverge, you always need to resort to raw SQL query, then it becomes a pain to refactor. The good ORM libraries relies on metaprogramming, then they're just weird SQL.
- groone 1y agoORM libraries have Value conversion functionality for such trivial examples https://learn.microsoft.com/en-us/ef/core/modeling/value-conversions https://learn.microsoft.com/en-us/ef/core/modeling/value-con...
- skydhash 1y agoNot really. It's all about the code you need to write. Instead of wrangling the data structures you get from the ORM which is usually similar to maps and array of maps. You have something that makes the domain logic cleaner and clear. Code for mapping data are simple, so you just pay the time price for writing them in exchange for having maintainable use case logic.
- galaxyLogic 1y agoTo me domain model is an Object-Oriented API through which I can interact with the data in the system. Another way to interact would be direct SQL-calls of course, but then users would need to know about how the data is represented in the database-schema. Whereas with an OOP API, API-methods return instances of several multiple model-classes. The way the different classes are associated with each other by method calls makes evidednt a kind of "theory" of our system, what kind of objects there are in the system what operations they can perform returning other types of objects as results and so on. So it looks much like a "theory" might in Ecological Biology, m ultiple species interacting with each other.
- Arch-TK 1y agoYou can model this "theory" in the database itself.
- nazgul17 1y agoMy understanding is, a database model is one that is fully normalized - design tables to have no redundant/repeated piece of information. You know, the one they teach you when you study relational DBs. In that model, you can navigate from anywhere to anywhere by following references. The domain model, at least from a DDD perspective, is different in at least a couple of ways: your domain classes expose business behaviours, and you can hide certain entities as such. For example, imagine an e-commerce application where you have to represent an order. In the DB model, you will have the `order` table as well as the `order_line` table, where each row of the latter references a row of the former. In your domain model, instead, you might decide to have a single Order class with order lines only accessed via methods and in the form of strings, or tuples, or whatever - just not with an entity. The Order class hides the existence of the order_line table. Plus, the Order class will have methods such as `markAsPaid()` etc, also hiding the implementation details of how you persist this type of information - an enum? a boolean? another table referencing rows of `order`? It does not matter to callers.
- setr 1y agoGenerally the ideal format for one problem is not the same as another. For example, to store a graph in a RDBMS, the ideal format is probably an adjacency list with a recursive query to iterate it. But in my app code, it’s probably easiest as an object-graph just pointing at each other. And in the context of my frontend, I don’t even want to talk about the graph, the user can only really talk about one node’s parent/child relationship at a time. There’s no one data model ideal for all scenarios — so why not have a different model for each scenario? Then I just need to figure out a way to transform between one model and the next, and whatever logic depending on that idealized data model can now be implemented fairly simply (since that’s the nature of a good data model - the rest of the logic often just falls out). So the data model you’re using then is localized to the domain/subject in question. You’re just transitioning the data between models as needed. A domain just being an arbitrary context — the persistence layer, or the UI logic, or even specific like I want my model for an accountant to reflect how an accountant UI page would organize it because I only understand 30% of what they’re asking me to do so keeping it “in their terms” makes things much easier to implement blindly. Or perhaps the primary purpose of this particular function is various aggregations for reporting, so I start off by organizing my dataset into a hierarchy that largely aligns with the aggregation groups. Once it’s aligned properly, the aggregation logic itself becomes utterly trivial to express You could even say that every time you query the database beyond a single table select *, you’re creating a new domain-specific data model. You’re just transforming from the original table representations to a new one. All domain modeling is specifically choosing a representation that best fits the logic you’re about to write, and then figuring out how to take the model you have and turn it into the model you want. Everything else on the subject is just implementation detail.
- sgarland 1y ago> For example, to store a graph in a RDBMS, the ideal format is probably an adjacency list with a recursive query to iterate it I know this was a minor point, but I think it speaks to the overall topic, so I'll poke at it. Adjacency lists are perhaps the worst way to store a graph / tree in RDBMS. They may be the easiest to understand, but they have some of the worst performance characteristics, especially if your RDBMS doesn't have Recursive CTEs. This starts to matter at a much lower scale than you might think; several million rows is enough to start showing slowdowns. This book [0] (Joe Celko's Trees and Hierarchies in SQL For Smarties) shows many other options, though it does lack the closure table approach [1], which is my preferred approach. And here, we come full circle back to the long-held friction between DBs and applications. You start mentioning triggers, and devs flinch, stating that they don't want logic in the DB. In every case I've ever seen, the replacements they come up with is incredibly convoluted and prone to errors, but hey, it's not in the DB. There is no reason to fear triggers, if and only if you treat them the same way that you'd treat code (because it is): added/modified/removed only via PRs, with careful review and testing. [0]: https://ia804505.us.archive.org/19/items/0411-pdf-celko-trees-and-hierarchies-in-sql-for-smarties-elsevier-2004/0411%20pdf%20Celko%20-%20Trees%20and%20Hierarchies%20in%20SQL%20for%20Smarties%20%28Elsevier%2C%202004%29.pdf https://ia804505.us.archive.org/19/items/0411-pdf-celko-tree... [1]: https://dirtsimple.org/2010/11/simplest-way-to-do-tree-based-queries.html https://dirtsimple.org/2010/11/simplest-way-to-do-tree-based...
- valenterry 1y ago> Should you start with a domain model from the get-go? Maybe, but it's risky because you may end up with domain objects that don't actually do a better job of representing the domain than whatever you have in your SQL tables. You absolutely should go with a domain model from the get-go. You can take some shortcuts if absolutely necessary, such as simply using a typealias like `type User = PostgresUser`. But you should definitely NOT use postgres-types inside all the rest of your code - that just asks for a terrible refactoring later. > It also feels awkward (and is hard to justify in a team) to map back and forth between domain model, sql SELECT row and JSON response body if they're pretty much the same, at least initially. Absolutely not. This is the most normal thing in the world. And, in fact, they won't be the same anyways. Don't you want to use at least decent calendar/datetime types and speaking names? Don't you want to at least structure things a bit? And you should really really use proper types for IDs. User(name: string, posts: string[]) is terrible. User(name: UserName, posts: PostId[]) is acceptable. So you will have to do some kind of mapping even in the vast majority of trivial cases.
- jiggawatts 1y agoAfter decades of experience, I'm starting to acquire a notion that most of modern web app development is simply the obstinate refusal to put the code where it really belongs: inside the database engine. Impedance mismatch, ORM, type generators, query parameterisation, async, etc... all stem from treating data as this "external" thing instead of the beating heart of the application. It terrifies me to say this, but sooner or later someone is going to cook up a JavaScript database engine that also has web capability, along with a native client-side cache component... and then it'll be curtains for traditional databases. Oh, the performance will be atrocious and grey-bearded wise old men will waggle their fingers in warning, but nobody will care. It'll be simple, consistent, integrated, and productive.
- NilMostChill 1y agoThat's....certainly a take. it hurts that it's not exactly wrong. but i don't think it's 100% right either, there are some things that you just can't do reliably, in current db engines at least. As soon as you start baking this kind of support in to the db all you have if a db engine that has all the other bits stuffed in it. They'll still have most of the issues you describe, it'll just be all in the "db layer" of the engine.