4 ms·
Was discussing your point with a colleague and the comment about time travel is very interesting. Could you give a little more detail the the point? With us, ev
by LukeEF 6y ago
Was discussing your point with a colleague and the comment about time travel is very interesting. Could you give a little more detail the the point? With us, everyone can be talking about a different time point in a branch whenever they like, and it has nothing to do with head. Two clients can look at different time points in our database and neither have to know about each other at all.
Also, for TerminusDB we don't use a multimaster coordination mechanism - we actually use the same sort of git approach.
- Fiahil 6y agoFirst, I need to say that I never tried TerminusDB, so I can't claim having a strong opinion on your approach :) Back to the time-travel. One of the most evident architecture when dealing with AI/ML/Optimisation is to design your application as a mesh of small, deterministic, steps (or scripts) reading input data and outputting results. As you would expect, output of one step is reusable by another one. Example: Script A is reading Sales data from source S, Weather data from source W; writing its result to A. Script B is reading data from source A, and Calendar from C; writing its result to B. In this example, we want to be able to do two things: 1) run a previous version of A with S and W from 2 weeks ago and assert the result it produced now is exactly identical to the one it produced at the time 2) run a _newer_ version of A with S and W from 2 weeks ago and compare its result from the one it previously produced. Of course, in the real-world, S, W, C, progress at different speed : new sales could be inserted by the minute, but the weather data would likely change by the day. So, you need a system that would allow you to read S@2fabbg and W@4c4490 while being in the same "repository". That's why git semantics are not a good fit: you need to have only one "branch" to ensure consistency and limit misunderstandings, but you want to "commit" datasets in the same repository at different pace. For that purpose, event sourcing is much better :) (BTW, git at its core, is basically event-sourcing) Kafka's architecture is actually the best solution.
- LukeEF 6y agoYou'll have to try :) Very interesting - thanks for the additional detail. Will have to think about how we might best represent in Terminus. We did a bunch of work for retailers in exactly the situation you describe.
- zachmu 6y agoDolt actually addresses this need exactly. You can query tables at two different revisions without checking out the two different branches. E.g.: `SELECT * FROM S AS OF '2fabbg';` `SELECT * FROM W AS OF '4c4490';` (Branch names work as well as commit hashes for the above). You can even do joins between the tables as they existed at those revisions: `SELECT * FROM S AS OF '2fabbg' JOIN W AS OF '4c4490' ON ...` As long as your data is actually relational, it's a pretty good fit.