5 ms·
Just to add some contrast to the mostly negative comments here (which have merit), this is interesting to me, not because it aims to hide the past, but because
by jonstaab 7y ago
Just to add some contrast to the mostly negative comments here (which have merit), this is interesting to me, not because it aims to hide the past, but because it makes time a first-class concept in the software development model, which most programming styles fail to do (e.g. most RDBMS frameworks add migrations on as an afterthought). I like this, and hope something like it catches on.
The problems with this approach seem solvable to me, albeit with more experimental magic that could explode:
- The resulting big ball of mud (and subsequent performance problems of a long pipeline of relations) could be compiled away, resulting in a single artifact. That is, you'd develop in append-only style, but when you "commit" your release, you'd end up with a single, optimized artifact for deployment, which would also be readable. This seems really nice to me, since your changes, while being based on the behavior of the system, would create a diff in the implementation of the system, potentially reaching way back into upstream events (in the example, the behaviors that block hot water and substitute cold water would just completely eliminate the first pane and simplify to adding cold water). This would let you see your changes from multiple perspectives. This approach also seems really friendly to fuzz-testing, which would give you a third look into the behavior of the system, and you could write tests based on the final state of the system after a number of given events.
- Migrating data structures actually seems easier to me for an event sourced approach, since you'd just re-project your domain models based on the new flow of events. b-threads would allow you to re-compile your event stream just like you re-compiled the source artifact (having parameterized events remains a problem, since your historical data could end up being incomplete and invalid based on new policy, you'd have to adopt a permissive schema to keep stuff that validation would otherwise reject).
I'll agree that b-trees don't really solve anything, but they do bring up some interesting questions that I think are worth asking. Datomic and Darklang I think are much more practical, and seem to dabble in the same sort of areas.
- hacker_9 7y ago>> when you "commit" your release, you'd end up with a single, optimized artifact for deployment I agree this would be nice, but think of this problem more deeply and you quickly run into major issues. The biggest one being that your compiler now has a large scale optimisation step to do completely automatically, better than a human. In order to make this work, you'd effectively have to build everything out of composable components, which you then can switch on/off with some sort of layering system. The problem here is how do you come to a decision on what those composable units look like? You don't know what you don't know - i.e. what those future layers you append are going to need to turn off or on. And coming back to optimisation - once you know the features you are combining together in an algorithm, only then do you know the data structures that fit the problem. Optimisation isn't a forward process when you turn parts off and end up with an efficient system. Optimisation requires a high level view, a low level view, an understanding of context, of memory requirements, of language capabilities, even leaps of logic. For the last part when I mean is that whilst a compiler may be able to reorganize something in the form a*b=c (and even then knowing when to do this is a book on it's own), it can't know that if I, for example, disable a feature that adds nested nodes to a tree structure, that it could rewrite functions to no longer require recursion and instead treat them like serial lists. You could rewrite the methods yourself and add them to your new layer, but you also will inevitably end up not fitting the previous composable structure. I think the deeper problem is related more to the tying of syntax to structure, as structure is ultimately your application's performance profile. Untying these two though could lead to a way forward for the append style paradigm.
- adamc 7y agoEven if you can compile away the performance problems, how do you escape the analytical mess that is left being? Patching software the old-fashioned way is expensive because we have to integrate the changes into the model -- essentially, rethink parts of the software. But the result is certainly likely to be easier to understand down the line. In a layered, "historical" model, I now have to understand the entire history, and correctly deduce how the history has changed the functioning of the original model. That strikes me as horrible.
- btown 7y agoIs this not a tooling problem? There are two compiler targets: the runtime, and the engineer. A system like this is sustainable iff it can compile a subset of the runtime-necessary information into a readable, interactive, sequential form. And this very much exists in the real world: every moddable video game, every audio/video tool accepting plugins, every multi-team business workflow, every browser plugin architecture, every SQL trigger... all are collections of independently developed state machines, intercepting a global event list and interrupting with their own events, just waiting to conflict with each other. Plugins and (blocking/yielding) extension points are the only real way to build software with massive feature surface areas. The tooling to visualize and debug these flows is IMO quite lacking, and I don’t think this is helped by systems engineers’ general love for all things textual. One needs to not only understand the flows they think will happen, but also fuzz the ones they don’t, and all this needs to be presented in a way that doesn’t overwhelm human working memory. I don’t think that’s a solved problem by any stretch. And I doubt the solution will look like our current text editors. But it’s something I think about quite often.
- yowlingcat 7y ago> The resulting big ball of mud (and subsequent performance problems of a long pipeline of relations) could be compiled away, resulting in a single artifact. So..something like `git rebase` or `git merge` for both code and data. That actually sounds cool, but certainly not trivial. Part of why I've always liked git, despite its warts, is that it's a robust tool that has never crashed on me, probably due to the immense amount of man-hours that have gone into hardening it to be usable for the kernel. According to my Google eng friends, there's even crazier levels of VCS sophistication going on to make the Google monorepo work. How hard would it be to make something like that for data? Call me crazy, but I've found it generally harder or maybe just _higher risk_ to reason about data, because incorrect assumptions result in bad things happening -- heisenbugs or zombie bugs. Given that we're talking about data migrations...what happens when you have a merge conflict? That's the hard part. Fuzz testing and writing tests based on the final state of the system? Also hard, but will be hard no matter what because making sure the end state doesn't contain undesired corner side effects could end up being a combinatoric explosion that results in data corruption and an alarmed user (or more). > "just re-project your domain models based on the new flow of events" just re-project your domain models? Sounds like a migration to me. I get a little annoyed by abstract models in the subject article because they 1) re-imagine things that already exist and 2) focus on making the easy things easier and not on making the hard things possible. For example, I even see this in your reasoning, which I don't agree with. Time is already a first class concept in most software development models. If you're using an imperative programming language, sequence is the default. Likewise for most RDBMS migration frameworks I've used. It's the data migration that's the problem. It's applying temporal evolution to data in a structured, testable way that scales and doesn't decay which is challenging. I've been thinking about writing a data migration framework for a while -- something like `alembic` but expressly for migrating data and testing that the migrated data is sane. Thing is, that's hard, and it gets harder the larger your data gets in volume and heterogeneity. At a certain problem, it's 100% a question of how well your understand your data, your UX, your business logic, and the history of your data. In fact, the majority of places I've worked at suffer from the situation where an old, inactive customer has bad data that no one bothers to fix because it's a poor usage of resources for the business. New ideas of programming "styles" or "models" that don't emerge from these pains -- are they going to solve them? Maybe. But this one doesn't seem to do so.
- spc476 7y agoI'm thinking about how this technique would apply to our codebase, and I'm not seeing it. The company does Caller Name ID (returns a name based on a phone number), and the code I work on is on the call path (as the call is being made). Our original code path for Android on CDMA went A, B, C, D and E. Then SIP came along, and we had to do A, B, E and D (C no longer required for SIP) but we still have to support CDMA. Then step B got more complicated, so we had to do B and B2 at the same time (basically, B and B2 are data queries internally), so CDMA: A, B & B2, C, D and E. SIP: A, B & B2, E, D. Then iPhone support comes along, and it doesn't matter if it's CDMA or SIP, we have to do A, B2, conditionally B (depends upon result of B2, so here we can't do both at the same time) and finally E2. And keep all the previous behavior intact. In our case, we aren't replacing old code, but keeping it around and adding new features into the mix. I think either way we'll still end up with a big ball of mud.