4 ms·
While not new, I find the discussion around immutable storage important and interesting. Datomic[1] comes to mind. There are many scenarios and applications, wh
by mtrn 10y ago
While not new, I find the discussion around immutable storage important and interesting. Datomic[1] comes to mind. There are many scenarios and applications, where it would be nice to just write things to a structured log and build an application on that log.
Unfortunately there are some roadblocks, like little information and conversation about the topic and no real mainstream implementation of these ideas.
I've been working on a system, that has is based on an immutable data layer and it is just a wonderful thing to work with, as long as you have enough disk space.
[1] http://www.datomic.com/ http://www.datomic.com/
- smallnamespace 10y agoI'm curious, how is this different from using an event oriented database, or using your normal DB but changing the schema to be oriented around immutable events?
- mamcx 10y agoNow I trying to use Event Sourcing (mainly for sync) for a "normal" invoice app, I find several problems. The thing is that the event log is good for write but awful for read. Re-compute everything is problematic and the true is 90% of the time you want the last version, not all the history. Also, the split in logic from the app / engine is stupid. I think is good if we move past the idea that the DB engine is more dumb than a rock, and embrace it. But current tools are not ideal (sadly, I know what is live like this: I have used Visual FoxPro!) Where I think this could lead to something great is if the DB is alike this: - Relational - You have normal tables. Consider them your up-to-date cache. Most scenarios will be covered here (ie: The log recomputing is NOT at the app level, but at the DB level). - You have a event log where everything is stored, plus your "tables" for the info you need up-to-date. But you don't need all to be a table if not make sense, and do full re-computation if desired. - Your backup, sync and load history is around the event log. - You don't even need a "table" if you wanna only a "index", this lead to me: - You submit a data request (like a POST), and the DB convert it to the tables and log, but: - You can instead (or also!) hook here and build secondary indexes and do other stuff. You can do it async, put the computation in a queue, do validations, etc as fit. So, what is not here and need tooling is marry the "normal tables" + "event log" + "logic that route this", so putting: Data -> WAL -> Router -> Event Log | Tables | Index | ping external tools And Request Data -> Router -> Get it from:Event Log | Tables | Index | ping external tools
- icebraining 10y agoHave you seen PipelineDB? It's Postgres, but with something called "continuous views", which are like regular views except they're cached and updated on writes, not re-calculated each time they're read. Sounds like one could use plain tables as the "event log", and these views as the up-to-date cache. And aggregation queries work too, so your regular table can be more than just a cache. I'm not affiliated, not even a user, it just sounded interesting when it was submitted here.
- mamcx 10y agoI have not use it, I my first glance make me think is only about analytics, but maybe that could work. I will try and see how well...
- Fergi 10y agoJeff from PipelineDB here. Yes, PipelineDB is an analytics product designed to continuously query high volumes of streaming data. The top use cases we see are building realtime reporting dashboards and realtime monitoring and alerting systems. PipelineDB excels at workloads where you know the queries you want to run ahead of time, where your workload fits within the confines of SQL, and where there is a high degree of distillation (aggregations, sliding windows, etc.). We have large customers including Charter Cable / Time Warner Cable, MediaMath, SmartNews, MOAT Analytics, Cradlepoint Networks, Paddy Power Betfair, and others using the system at scale and we offer a clustered edition of PipelineDB under a commercial license, but PipelineDB's single server edition is open-source. We have a live chat room here for technical questions: https://gitter.im/pipelinedb/pipelinedb https://gitter.im/pipelinedb/pipelinedb
- yrashk 10y agoYou are quite right, event sourcing on its own is not great for read, and the "re-compute everything" (aka foldl) is problematic. CQRS solves this by producing a read-side that aggregates events into your domain. Classic event sourcing also attempts to solve the re-computation problem by snapshotting. It is great that you mention that 90% of the time you want "the last version". That was one of my discoveries with event sourcing as well. So what I came to is that re-computation (foldl) is a nice generalization, but isn't very practical. Instead, index the events and build the state on demand by querying the indices. This way, [plurality enabling] domain modelling moves to the code (as opposed to the database) and allows for quick changes in how you see the information. You can read more about it at https://blog.eventsourcing.com/why-use-eventsourcing-database-6b5e2ac61848#.o1q5yb3p9 https://blog.eventsourcing.com/why-use-eventsourcing-databas... and https://blog.eventsourcing.com/lazy-event-sourcing-ed7e59007e17#.v5fa8aqbt https://blog.eventsourcing.com/lazy-event-sourcing-ed7e59007... as well as other articles on that blog.
- clusmore 10y agoOfftopic to the OP but related to Datomic. I saw Rich Hickey's talk about Datomic recently and immediately thought it would be ideal as part of a distributed runtime. Without having actually programmed in Erlang, but having seen plenty of Joe Armstrong's talks, I understand that the point of replicating data in messages rather than passing around pointers is so that if a machine dies, the data isn't lost. But if you built a distributed runtime that used Datomic as a distributed heap on which you allocate large amounts of data which you expect to pass around a lot, you could instead pass around only pointers to the data and not worry about data loss, and save big time on message throughput. Does anybody know if something like this already exists?