21 ms·
Revisiting the principles of data-oriented programming
- readthenotes1 4y agoI haven't loved debugging systems whose primary data structure is a Map.
- qsort 4y agoI think it very much depends on what problems you're trying to solve and whether or not proper data types have been defined. If your primary data structure is Map<Integer, List<String>> we have a huge problem. On the other hand, if your primary data structure is Map<CustomerId, List<Purchase>> Then I'd rather see that than IPurchaseMappingByCustomerIdAbstractFactory or whatever other abomination OO priests will conjure. Generally speaking, generic structures are simpler and they allow for easier transformations.
- marcosdumay 4y agoThe article doesn't explain, but links into a deeper one. The author really means you should use Map<Integer, List<String>>. He also seems to be unaware that you can have generic code with specific types.
- weavejester 4y agoYou can still use types, as long as they don't encapsulate the data. Admittedly this is hard to do in some languages.
- skippyboxedhero 4y agoI am not entirely sure that most data-oriented programs do go this way. You can split functionality out of the data but the data can still be represented with an object or whatever. I would agree though, you might as well be using Python at that point. An example of what I mean is Spring. Obviously, from what I recall, that goes to other extreme with lots of XML configuration. But there is no need for vague types that can cause all sorts of mischief at runtime. The key idea is splitting code from data, not necessarily the representation of the data (although that can come into it if you have performance-sensitive apps).
- mtVessel 4y agoThere's something called "Data-Oriented Programming" and something else called "Data-Oriented Design". I can never remember which is which. This post changes nothing.
- GolDDranks 4y agoI think the abbreviated forms of the paradigm are often more "stable" than the full names, because people keep hand-waving the names and thus mixing them up. For me, "DOD" is the thing where you are very performance-oriented and you have flat, cache friendly arrays of data and affinity to ECS (entity-component-system) stuff etc. This is clearly not it, but eyeballing it, the ideas seem somewhat compatible with it. (Except the immutability part.)
- ArrayBoundCheck 4y agoOne is ridiculous and uses immutable data. The other is made famous by a guy in a Hawaiin shirt and preforms well. Designer on vacation might help you to remember which is which
- Tomis02 4y agoIndeed. The guy in Hawaiian shirt writes simple, fast code with the minimum amount of abstractions, and therefore can afford to finish work early and go to the beach. Everyone else is working overtimes untangling a mess of objects, principles and hierarchies.
- mistrial9 4y agoLarry Wall is that you!?
- peoplefromibiza 4y agoMost likely he's Mike Acton.
- javajosh 4y agoTwo concepts that seem entirely missing are "ownership" and "transformation". The root of ownership is "system of record" with intermediate systems combining data and doing transformation. "Getting data" becomes a question of where, in a differentiating tree that has it's root in the SoR, you connect to, and the trade-offs that implies. This post (and the book it points to) is perhaps teaching a new generation what has been known for a long time: the "body" of your business is the data, not the code. E.g. if you have limited space on a thumbdrive and can only keep one thing in a datacenter fire, your database or your codebase, you keep the database.
- ncmncm 4y ago"Anything"-oriented programming is dumb. Every big problem is a collection of smaller, different problems. Different problems call for different approaches. Sometimes the best approach has data-oriented features, sometimes object, sometimes functional, sometimes piped. For big problems you want a language good at all of them. It is why C++ continues to grow. Some complain about that, but every single feature got there over fierce opposition by making some common programming problem more tractable.
- irrational 4y agoI think this is also a reason people try to use JavaScript for so many things. As a multi-paradigm language you can do OOP, functional, etc. in it.
- sidlls 4y agoPeople try to use JavaScript for so many things because it's (surface-level) inexpensive, not because it's good at anything.
- galaxyLogic 4y agoIt is good enough
- sidlls 4y agoYes, like PHP. In so many ways.
- eurasiantiger 4y agoMind-reading and zealotry, you must be a priest. What is your denomination?
- drpyser22 4y agoIts a good thing that these paradigm exist so you know their value, and eventually understand when they're appropriate. You're right that there may not be a one-size-fits-all.
- xixixao 4y ago> Adherence to this principle in OOP means aggregating the code as methods of a static class. This is not OOP, this is a way to do functional programming in a class-based language that lacks top-level function declarations / modules. While this might seem a nit pick it makes me sceptical about the rest of the content.
- osigurdson 4y agoI think possibly what they meant is "when using an object oriented language like Java or C#", aggregate the code as methods of a static class.
- hinkley 4y agoThat's not necessary either. What we're talking about here is an analog for the real problem, which is "when using an object oriented language, write pure functions even though the language doesn't make you do it". Static functions introduce friction against making impure functions, but not an overwhelming amount of it. If it's all you have, then there are worse safety blankets, but it's not going to solve your lack of buy-in problem. If everyone is on board then you don't need static functions. If they aren't, static functions aren't going to save you, they're just going to extend the amount of time you suffer before you wise up and get a new job somewhere else.
- osigurdson 4y agoIf all of the data structures are maps of maps of maps, it seems that it would be obvious to use static methods in this situation. I don't understand the use case for instance methods with this arrangement.
- hinkley 4y agoLazy evaluation for one. Organization for another.
- alphanumeric0 4y agoI'm having a hard time thinking of a way code can ever be fully decoupled from data. When we decide it's better to have a name field rather than firstName and lastName does that mean we simplify NameCalculation.fullName to just return data.name? This seems to suggest we still have coupled code to data (the data structure being an object), it's just now a coupled function, but you have decoupled it enough to use NameCalculation in different contexts. Single responsibility classes are already recommended for reuse like this in OO. Also when it comes to data validation, OO performs all of this validation too and in a much more compact and code-oriented, extensible way. Why would I write a separate schema when the object itself knows what it will accept, what is optional, and what range the values should be in? I'd imagine the schema and code could become incoherent.
- drpyser22 4y agoFor validation, this approach would have you write a set of functions to validate the properties of the data. Nothing forbids a function that applies validation to inputs before returning a data object? Extensibility can be done through functional means(e.g. higher order functions, function composition, lens) or oop(strategy pattern and equivalents, code object composition and inheritance,...). Not sure what you mean by more compact and code-oriented? Code is always coupled to an interface, implicitly or explicitly. In the case of oop, code is coupled to the class, which can represent something specific with very concrete semantics (e.g. employee, author) or something generic that is meant to be subclassed(e.g. person).
- snidane 4y ago> Why would I write a separate schema when the object itself knows what it will accept, what is optional, and what range the values should be in? Because data is just data and the meaning to it is given at the time of application. If you want to couple validation to the data itself - how do you decide which N of the meanings to validate against?
- galaxyLogic 4y agoNo, data must have meaning, else it is meaningless. If you want to process the data you must use some language to access parts of the data. Data must have a symbolic representation, not juts be 1s and zeros. Or it can be a bit-stream, but even then you need a language that knows the different between 1 and 0. person.name extracts the field 'name' from the data. To manipulate the data, the program must know that there is such a field as 'name' it can ask for.
- andreareina 4y ago> Breaking this principle in FP means hiding state in the lexical scope of a function. If that's happening a lot that's not really FP anymore, is it?
- throwaway17_17 4y agoNothing about keeping values in functions is ‘non-functional’. That’s like saying that hard coding the quadratic formula inside some function instead of using a lambda as an input is ‘non-functional’. His language in that statement is poorly chosen. I am almost certain he is not implying any action at a distance ‘state’, he is trying to talk about including context for some data inside functions that operate on the data. It would be like hardcoding an ISBN-to-title list inside a function that takes a list of authors and there books as input for processing. I think he’s saying the ISBN-to-title should be a part of the data structure, and storing it inside the functions breaks these rules he has invinted.
- osigurdson 4y ago>> static boolean isProlific (Map<String, Object> data) { >> return (int)data.get("books") > 100; >> } Could anything be more confusing with a large code base? Also, lots of nice key not found and invalid cast exception errors to debug with this approach. Sometimes boxing makes a material difference to performance as well.
- osigurdson 4y ago>> Take, for example, AuthorData, a class that represents an author entity made of three fields: firstName, lastName, and books. Suppose that you want to add a field called fullName with the full name of the author. If we fail to adhere to Principle #2, a new class AuthorDataWithFullName must be defined Wait, what? Just add another field/property to the existing class. It's a silly example anyway as normally you would just add a function to concatenate the two strings. The stated advantage is to be able to add the new property "on the fly". I suppose this means without changing the code. It does beg the question "what can existing code possibly do with this" (other than display it in a generic way or count the number of fields)? Furthermore, adding something new is rarely much of a problem as it is a non-breaking change. A more difficult example would be removing the "firstName" field. Assessing the impact of such a change in a large code base would be extremely difficult. Get good a grep and hope that the test suite is comprehensive.
- bo0O0od 4y agoAwful stuff. These principles could only ever make sense in a dynamic language since it's mostly manually enforcing some of the basic functionality of a type system, but the fact that he also tries to argue this style could be used in a language like C# throws that defense out the window. https://blog.klipse.tech/databook/2022/06/22/generic-data-structures.html https://blog.klipse.tech/databook/2022/06/22/generic-data-st... The examples also contradict his other principles, i.e. immutability.
- weavejester 4y agoIt's not impossible to type check heterogeneous maps at compile time, but most static type systems don't support this. I think you'd certainly see much more friction trying to program like this in C# than you would in Clojure.
- bo0O0od 4y agoI agree, but I also think if they author either knew what they were talking about or wasn't just trying to sell more books they'd make that clear. Rather than just trying make the case for these patterns in languages that don't suit them.
- zasdffaa 4y ago> It's not impossible to type check heterogeneous maps at compile time, but most static type systems don't support this I guess you mean dependent types[1], but if you don't, I'd appreciate an elaboration. If you do mean DTs, how might it look for a hetero collection? [1] If anybody has any good intros to dependent typing in C#, that'd be much appreciated. A web search throws up some pretty intimidating stuff.
- siknad 4y agoDependent types are types that depend on values, possibly runtime values. In C# types can only depend on other types when using generics: List<T> depends on T. In C++ there is std::array<T, n> (array with length encoded in type), where n must be known in compile time. With full dependent types one can write generic types like std::array and use them with runtime parameters. In dependently-typed languages there are two main types: Sigma (dependent pair) and Pi (dependent function). Example of Sigma type in pseudo C#: (uint n, Array<int, n> a); // array of any size Pi: Array<T, (n + 1)> push(Type T, uint n, Array<T, n> a, T x) { ... } Array<int, 3> x = push(int, 2, [1, 2], 3); // [1, 2, 3] Generic function f<T> is similar to a dependently typed one with `Type T` argument (requires first class types and many DT-langs have them). Values, on which types may depend, shouldn't be mutable, while C# function arguments are mutable. A bit larger example: void Console.WriteLine(string format, params object[] args); Using dependent types you can transform `format` into a heterogeneous array type containing only arguments specified in the format string. WriteLine("{%int} {%bool}", ?); // ?'s type is an array with an int value and a boolean value. A heterogeneous map may be implemented as a map T with keys mapped to types and another map where keys are mapped to the values of a corresponding type in T. Probably this is not a good representation, but it is a valid one.
- banq 4y agoit actually is Domain Driven Design, domain data-oriented programming
- g9yuayon 4y agoReading the discussion here, I can't help but thinking that people are defending their own philosophies: OOP vs FP vs DOP vs etc. I wish the author had killer applications or killer examples in different categories, like can I code an operating system easier, can I code a database easier, can I create a complex streaming job easier, can I write a library as complex as Apache BEAM easier, can I write a compiler easier, can I create a web framework easier, can I write a JSON parser easier, you get the idea. Or maybe examples that contrast existing solutions: how do I use DOP to write a better RxJava? how do I use DOP to write a better SqlLite? How do I use DOP to write a better graph library? How do I use DOP to write a better tensor library? how to do use DOP to write a better Time/Date library? You know, something that's so compelling and so obvious.
- throwaway894345 4y agoYou'll have to define "OOP" first. Everyone thinks its defined, but even among OOP proponents there isn't consensus ("it's about message passing", "it's about encapsulation", "it's about inheritance", "it's about dot-method syntax", etc).
- deltaonefour 4y agoOne way to define it is to look at the thing that's unique to OOP that isn't used in any other paradigm. In OOP data and method are unionized into primitives that form the basic building blocks of your program. This is unique to OOP. It must be the definition. Defining it in terms of message passing, encapsulation and inheritance are not as good because these concepts are used in other paradigms as well.
- danielscrubs 4y agoWouldn’t closures and let’s say “type-driven” also support your definition? It’s quite tricky as sometuple.fun() and fun(sometuple) might as well be interchangeable. If you really want to support a formal definition it would probably be best to have it in denotational semantics… not an easy task.
- eurasiantiger 4y agoImmutability is nice but now your data migration needs quadratic memory
- forgotusername6 4y agoRedux seems to have points 1-3. There is typically no schema though.
- brunooliv 4y agoTruly misguided article as most of the things by the author, unfortunately. I was so excited about the book "Data-oriented programming" when it was first being released...it was so heavily publicized as well that it was constantly in my face which likely pushed me over the edge to give it a shot and buy it. Unfortunately, not all that glitters is gold. It feels extremely beginner oriented, only touched basic concepts taught at uni level and it shows a huge disconnect between the theory and the real world work of a developer leveraging data in any way, shape or form. Don't buy the book, it's so not worth it.
- blain_the_train 4y agohow did it show a disconnect?
- brunooliv 4y agoEssentially, by pushing the discussed topics as the One True Way and almost "purposedly" choosing to not discuss trade-offs and talking about alternatives or disadvantages. In the real world, trade-offs are very important
- pca006132 4y agoI thought this is talking about data-oriented design, which focuses on the data layout to make programs more efficient, e.g. structure of array that can be more cache friendly in some cases. > Principle #2: Representing data with generic data structures. OK probably this is not what I expected.
- revskill 4y agoNo code example ? No, it's a sign of bad book itself. One code example is worth 1000 images, and 1 image is worth 1000 words. Always use code block to illustrate your point, as it help reader understand better your point. Writing a book is more about getting reader into your thought rather than make them think.
- bob1029 4y agoI believe the easiest way to think about it is to get away from your programming tools and to start modeling your problem domain as tables in excel. Once you have a relational schema that the business can look at and understand, then you go implement it with whatever tools and techniques you see fit. This is what “data-oriented” programming means to me. It’s not some elegant code abstraction. It’s mostly just a process involving people and business. Even for non serious business, these techniques can wrangle complexity that would otherwise be insurmountable. I still think the central piece of magic is embracing a relational model. This allows for things like circular dependencies to be modeled exactly as they are in reality.
- jackosdev 4y agoThis confused me as when I hear data-orientated I think data structures that are optimised around minimising CPU cache misses by using better alignments, using enums where possible, not storing results of simple calculations etc. There is a popular book and popular talks on the subject. Probably confuses other people as well I'd imagine
- bob1029 4y ago> I think data structures that are optimised around minimising CPU cache misses by using better alignments, using enums where possible, not storing results of simple calculations etc. You might be surprised to learn that modeling your problem in terms of normalized relational tables ultimately achieves similar objectives. The more normalized, the more packed you will find their in-memory representations.
- viebel 4y agoDOP is not the same as DOD [1] 1: https://blog.klipse.tech/visualization/2021/02/16/data-related-paradigms.html https://blog.klipse.tech/visualization/2021/02/16/data-relat...
- frogulis 4y agoRich Hickey's talks "Simple Made Easy" [1] and "Effective Programs" [2] provide a better explanation of these ideas IMO. The specific definition of "simple" is pretty crucial. [1] https://youtu.be/LKtk3HCgTa8 https://youtu.be/LKtk3HCgTa8 [2] https://youtu.be/2V1FtfBDsLU https://youtu.be/2V1FtfBDsLU
- jokoon 4y agoI love DOP, and I loathe OOP. But when you use a framework that enforces OOP, it's quickly difficult to use DOP.