6 ms·
Data, objects, and how we're railroaded into poor design (2018)
- evrennetwork 1y ago[dead]
- chuzz 1y agoGood post, for what is worth Java is slowly and painfully correcting course with features like records and project Valhalla. As with the other language improvements though we will have to live with the tech debt for decades to come…
- Timwi 1y agoWhile C# certainly has tech debt too, the specific feature the author mentions (value types) has been there since version 1. C# pioneered async/await in version 5, and has a plethora of functional features like generics, type inference and pattern matching now. Java is lagging behind like a wounded cow...
- lmm 1y agoI've had thoughts along this line for a while. I think Scala does better than this article gives credit for; case classes are a significant step in the right direction, particularly post-Scala 3 (or with Shapeless in Scala 2) where you have many tools available to treat them as records, and you can distinguish practically between case classes (values) and objects with identity even if in theory they're only syntax sugar. It also offers an Erlang-style actor system if you want one. In my dream language I'd push this further; case classes should not offer any object identity APIs (essentially the whole of java.lang.Object) and not be allowed to contain non-value classes or mutable fields, and maybe objects should be a bit more decoupled from their state. But for now I wouldn't let perfect be the enemy of good.
- arethuza 1y ago"objects should be a bit more decoupled from their state" Do you mean allowing the "class" of an object to be changed - CLOS can do that. Mind you it's a long time since I wrote any code using CLOS and even then I'm pretty sure I never used change-class.
- lmm 1y agoThe thing I'm envisioning is something akin to typeclass instances / trait impls, but specialised to the case where you have a service with identity rather than being for general function implementation. Just making the bridge between the "bag of state" piece and the "interface implementation accessible via a name/reference" piece a bit more of a first-class citizen.
- ratmeadow 1y agoIMO the nice thing about Erlang and Elixir is their foundation of representing data is rock solid. Because data is fully immutable, you get a lot of nice things "for free" (no shared state, reliable serialisation, etc). And then on top of that you can add your interfacey, mutable-ish design with processes, if you want. But you will never have oddities or edge cases with the underlying data. In contrast with languages like C++ and Java where things are shakey from the ground up. If you can't get an integer type right (looking at you, boxing, or you, implicit type conversions), the rest of the language will always be compensating. It's another layer of annoyances to deal with. You'll be having a nice day coding and then be forced to remember that int is different to Integer and have to change your design for no good reason. Perhaps you disagree with Erlang's approach, but at least it's solid and thought-out. I'd take that over the C++ or Java mess in most cases.
- jasode 1y ago>IMO the nice thing about Erlang and Elixir is their foundation of representing data is rock solid. Because data is fully immutable, you get a lot of nice things "for free" (no shared state, reliable serialisation, etc). [...] >In contrast with languages like C++ and Java where things are shakey from the ground up. Yes, immutable does provide some guarantees for "free" to prevent some types of bugs but offering it also has "costs". It's a tradeoff. Mutating in place is useful for highest performance. E.g. C/C++, assembler language "MOV" instruction, etc. That's why performance critical loops in high-speed stock trading, video games, machine learning backpropagation, etc all depend on mutating variables in place. That is a good justification for why Erlang BEAM itself is written in C Language with loops that mutate variables everywhere. E.g.: https://github.com/erlang/otp/blob/master/erts/emulator/beam/erl_thr_queue.c#L336 https://github.com/erlang/otp/blob/master/erts/emulator/beam... There's no need to re-write BEAM in an immutable language. Mutable data helps performance but it also has "costs" with unwanted bugs from data races in multi-threaded programs, etc. Mutable design has tradeoffs like immutable has tradeoffs. One can "optimize" immutable data structures to reduce the performance penalty with behind-the-scenes data-sharing, etc. (Oft-cited book: https://www.amazon.com/Purely-Functional-Data-Structures-Okasaki/dp/0521663504 https://www.amazon.com/Purely-Functional-Data-Structures-Oka...) But those optimizations will still not match the absolute highest ceiling of performance with C/C++/asm mutation if the fastest speed with the least amount of cpu is what you need.
- high_na_euv 1y agoI'd say majority of programming languages struggle with elegant, robust and exhaustive check error handling
- danieltanfh95 1y agoclojure exists.
- cbsmith 1y agoI didn't like how this essay misunderstood the design principles in Java's class files. With dynamic binding to the runtime, you can't know for certain the layout of a data structure in memory (e.g. what is the ideal memory alignment?). If the class file is untrusted, you can't even be sure you have a valid data structure in the first place. So allocating the array and then assigning elements one at a time is what you do.
- imtringued 1y agoIt also appears to be a nonsensical complaint in general. Why abuse .class files for data storage when you can choose any data format you like? I've only seen this pattern used in contexts, where there is no filesystem to begin with.
- Timwi 1y ago> Why abuse .class files for data storage The only alternative is to have the data in a separate file, which needs to be available, read in, and parsed. .NET/CLR provides a mechanism to bake large objects into the assembly and I don't see that as abusive. It's way more convenient when you can treat the object as being “just there”.
- gf000 1y agoWell, jars can have resource files included - I would say this is a more-than-solved problem.
- nopurpose 1y agoSo much written about relation between objects and data, but not a single mention of Lisp and derivatives?
- mort96 1y agoAn excellent opportunity for you to elaborate on the connection, since I'm not seeing it.
- Timwi 1y agoSame. Lisp’s selling point is that “code is data” — not objects.
- deterministic 1y agoAll code is data. Many languages (Haskell for example) can directly manipulate code as data (macros). The unique thing about lisp is that the code is represented as a car/cons list. Other languages could do the same when writing macros. However most have chosen not to.
- deterministic 1y agoLisp doesn’t have a monopoly on “data”. And most Lisps are not functional (setq/setf). Closure is different of course. But not more functional than Haskell for example.
- DrBeryl 1y agoHi, Ruby seems to take the opposite approach, but does it still resolve the problems exposed in this article? I'm a beginner in development, knowing the basics of a few languages (mostly Python and JS). Being very sensitive to design and logic, I want to choose the right path forward. I want to learn a language that makes sense and help me to think rigorously.
- 2d8a875f-39a2-4 1y agoMaybe it's just me but I find the complaint confusing and the suggested remedy absent in TFA, despite reading it twice. Data comes from outside your application code. Your algorithms operate on the data. A complaint like "There isn’t (yet?) a format for just any kind of data in .class files" is bizarre. Maybe my problem is with his hijacking of the terms 'data' and 'object' to mean specific types of data structures that he wants to discuss. "There is no sensible way to represent tree-like data in that [RDBMS] environment" - there is endless literature covering storing data structures in relational schemas. The complaint seems to be to just be "it's complicated". Calling a JSON payload "actual data" but a SOAP payload somehow not is odd. Again the complaint seems to be "SOAP is hard because schemas and ws-security". Statements like "I don’t think we have any actually good programming languages" don't lend much credibility and are the sort of thing I last heard in first year programming labs. I'm very much about "Smart data structures and dumb code works a lot better than the other way around" and I think the author is starting there too, but I guess he's just gone off in a different direction to me.
- arethuza 1y agoMy main complaint with SOAP was how leaky an abstraction it inevitably is/was - it can look seductively easy to use but them something goes wrong or you tried to use it across different tech stacks and then debugging would become a nightmare. When I first encountered RESTful web services using JSON the ability to easily invoke them using curl was such a relief... (and yes, like lots of people, I went through a phase about being dogmatic about what REST actually is, HATEOAS and all the rest - but I've got over that years ago). NB I also am puzzled as to the definition of "data" used in the article.
- BlackFly 1y agoYeah, a leaky abstraction with abstraction inversion on top of it! So within the actual payload there was a method identifier so you had sub-resource routing on top of the URL routing just so you could have middleware handling things on the payload instead of in the headers... So you had an application protocol (SOAP) on top of an application protocol (HTTP).
- 1y ago
- dzonga 1y agoclojure + lisps have this approach - everything is data. I encourage everyone to dabble in clojure. it changes your mindset. that when you go back to your regular language you will write better programs.
- DrBeryl 1y agoApparently, Ruby takes the opposite approach: everything is object. I bet it leads to a completely different way to think and design, but at least it does not mix the two, is it right? Do you recommend it? (I have a choice to make between Ruby and JavaScript.)
- mananaysiempre 1y ago> Extensibility > [Data] The schema gives us a fixed set of variants, over which you can of course write any function you want. > [Objects] We have a fixed set of exposed operations, but different variants can be constructed (including an improved ability to evolve a variant without impacting clients). No mention of the expression problem? The TL;DR is, sometimes[1] we want both. And sometimes[2] it’s an exceptionally good idea for a lot of slightly different sets of variants to coexist withn the same program, which there also isn’t really a satisfactory solution for. [1] https://www.craftinginterpreters.stuffwithstuff.com/representing-code.html#the-expression-problem https://www.craftinginterpreters.stuffwithstuff.com/represen... [2] https://nanopass.org/ https://nanopass.org/
- emorning4 1y agoWhy is the most relevant comment at the bottom of the page? The expression problem absolutely captures the nature of the issue and a lot has been written about it.
- Timwi 1y agoNo mention of C#, which of all the languages I've seen makes the best distinction for what he wants: structs/“value types” = data, classes/“reference types” = objects. He briefly mentions that Java is in the process of adding that; C# has had it since version 1. Now, 15 versions later, C# has acquired all the modern features from all paradigms of programming, including type inference, record types, and an incredibly full-featured pattern matching system. Java is laughably lagging behind. He also praises async/await as a revolutionary leap forward for programming languages, but without a mention of the language that invented/pioneered it (that's right, it's C#, in version 5). Now, I'm not going to sit here and claim that C# is perfect for all purposes. It has made mistakes which have to live on due to backwards compatibility. The mutability of structs is, in my opinion, one such mistake. But even so, of all programming languages I've seen, C# fits this author’s ideas like hand in glove. Edit: the scenario he describes of putting a large array or other static data structure directly into the code is also better supported by modern C# compilers. For small structures you still get the element-by-element-adding bytecode, but there's also a serialization format that can bake large objects directly into the assembly.
- gf000 1y agoI do agree that the article could be summarized as "value types vs reference types, a language should have both". But your idea of Java is pretty dated - it has type inference, full algebraic types (records and sealed classes), pattern matching (switch expression - though many of its more advanced features are TBD/experimental yet). And then Java has virtual threads, which C# suddenly wants to add as well, that language is definitely on the same road as C++, and this "adding every conceivable feature" is not sustainable, it will crumble under its own way. Java is much more conscious of it, which I greatly value.
- alphazard 1y agoPart of the mismatch here is that people discover programming languages as tools for thinking, maybe better than any other thinking tool they've ever encountered, and with the unique ability to verify one's thinking. Then they want to do all their thinking with this tool, including all of their design. IME this leads you to consider more and more featureful languages which are worse and worse at actually building software, and never match the flexibility of a tool like pen and paper. Poor design is your own fault. Write more, draw more, prototype more. You may need to develop your own notation, you may need to get better at drawing or invest in a drawing tool. You may need to learn another programming language, which you use only for prototyping. The design of your system does not need to be perfectly represented in your source code. The source code needs to produce runnable machine code, which behaves in the ways that your design dictates, and that's the only link between the design, the code, and the running system. Programming languages today are pretty good at producing working software, but not very good at doing that and designing systems and communicating designs and documenting choices, etc.
- kmac_ 1y agoWell said. Languages are for implementation, not design. There isn't a single language that connects both. Also, every popular language has multiple patterns that communicate code intentions. We can add to that conventions added by frameworks, companies, and even teams.
- jact 1y agoOCaml’s object system is unfairly maligned by its users. It’s unfortunate because object types allow for some of the most useful kinds of polymorphism and people don’t reach for them often enough. However, OCaml’s first class modules are frequently a useful and serviceable alternative. I think they’re a precise middle ground between what we here call “objects” and “data.”
- etbebl 1y agoThe author dismisses C++ out of hand, but I really think it does a pretty good job of making this a non-issue. Want a structured value type? Sure, that's a struct with public fields by default, passed by value with automatic copy constructor and assignment functions. Want a mutable type that's encapsulated and needs to do something special to be cloned? Sure, that's a class passed by unique_pointer or reference, with non-default (or deleted) copy constructor and assignment functions and private fields by default. Every language I've used since then feels like it makes this issue needlessly complicated and implicit.
- deterministic 1y agoC++ developer here (30+ years). C++ is really missing support for sum types. The Haskell JSON example shows how useful it would be to have native support for it. Yes you can build your own but it’s pages of boiler plate code.