11 ms·
Dolt is Git for Data: a SQL database that you can fork, clone, branch, merge
- bitslayer 6y agoIs it for versions of the database design or versions of the data?
- deleted 6y ago[deleted]
- skybrian 6y agoBoth. Schema changes are versioned like everything else. But depending on what the change is, it might make merges difficult. (I haven’t used it; I just read the blog.)
- fiedzia 6y agoBTW. I wish all databases versioned their schema and kept full history. This should be a standard feature.
- codesnik 6y agokeep history just as a changelog to see when and what was changed? or something with versions you actively can revert to? I suppose one of the problems is that changing schema is usually interleaved with some sorts of data migrations and conversion and those I have no idea how to track without using some general scripting language, like migrations in many frameworks which are already there.
- fiedzia 6y agoJust storing a changelog would be very useful. I could for example compare db version some app was developed for with current state. Most companies will store schema migrations in git, but doing anything with this information is difficult. Having this in db would make automated checks easier.
- strogonoff 6y agoYou can also use Git for data! It’s a bit slower, but smart use of partial/shallow clones can address performance degradation on large repositories over time. You just need to take care of the transformation between “physical” trees/blobs and “logical” objects in your dataset (which may not have 1:1 mapping, as having physical layer more granular reduces likelihood of merge conflicts). I’m also following Pijul, which seems very promising in regards to versioning data—I believe they might introduce primitives allowing to operate on changes in actual data structures rather than between lines in files, like with Git. Add to that sound theory of patches, and that’s a definite win over Git (or Doit for that matter, which seems to be same old Git but for SQL).
- teej 6y agoThe fact that I can use git for data if I carefully avoid all the footguns is exactly why I don’t use git for data.
- pradn 6y agoGit is too complicated. It's barely usable for daily tasks. Look at how many people have to Google for basic things like uncommitting a commit, or cleaning your local repo to mirror a remote one. Complexity is a liability. Mercurial has a nicer interface. And now I see the real simplicity of non-distributed source control systems. I have never actually needed to work in a distributed manner, just client-server. I have never sent a patch to another dev to patch into their local repo or whatnot. All this complexity seems like a solution chasing after a problem - at least for most developers. What works for Linux isn't necessary for most teams.
- Klwohu 6y agoProblematic name, could become a millstone on the neck of the developer far into the future.
- twobitshifter 6y agoThis is cool, but the parent dolthub project is even cooler! Dolthub.com
- justincormack 6y agoI collected all the git for data open source projects I could find a few months back, there have been a bunch of interesting approaches https://docs.google.com/spreadsheets/d/1jGQY_wjj7dYVne6toyzmU7Ni0tfm-fUEmdh7Nw_ZH0k/edit?ts=5fc6a2a5#gid=0 https://docs.google.com/spreadsheets/d/1jGQY_wjj7dYVne6toyzm...
- chub500 6y agoI've had a fairly long-term side project working on git for chronological data (data is a cause and effect DAG), know of anybody doing that?
- michaelmure 6y agoIt might not be exactly what you are looking for, but git-bug[1] is encoding data into regular git objects, with merges and conflict resolution. I'm mentioning this because the hard part is providing an ordering of events. Once you have that you can store and recreate whatever state you want. This branch[2] I'm almost done with remove the purely linear branch constraint and allow to use full DAGs (that is, concurrent edition) and still provide a good ordering. [1]: https://github.com/MichaelMure/git-bug https://github.com/MichaelMure/git-bug [2]: https://github.com/MichaelMure/git-bug/pull/532 https://github.com/MichaelMure/git-bug/pull/532
- glogla 6y agoThis one seems to be missing: https://projectnessie.org/ https://projectnessie.org/
- justinclift 6y agoIt's also missing DBHub.io. ;)
- LukeEF 6y agogreat list - i talked through it and what each tool does in a MLOps Community meetup: https://www.youtube.com/watch?v=r5uxntl_hWg https://www.youtube.com/watch?v=r5uxntl_hWg
- andrewmcwatters 6y agoReminds me a bit of datahub.io, but potentially more useful.
- pizzabearman 6y agoIs this mysql only?
- zachmu 6y agoIt uses the mysql SQL dialect for queries. But it's its own database.
- iamwil 6y agowhich db does it use?
- zachmu 6y agoIt is a database. It implements the MySQL dialect and binary protocol, but it isn't MySQL. Totally separate storage engine and implementation.
- jrumbut 6y agoIt's amazing this isn't a standard feature. The database world seems to have focused on large, high volume, globally distributed databases. Presumably you would't version clickstream or IoT sensor data. Features like this that are only feasible below a certain scale are underdeveloped and I think there's opportunity there.
- fiddlerwoaroof 6y agoDatomic has some sort of zero-cost forming of the database: it’s “add-only” design makes this cheap.
- qbasic_forever 6y agoEvery DB engine used at scale has a concept of snapshots and backups. This just looks like someone making a git-like porcelain for the same kind of DB management constructs.
- zachmu 6y agoIt's not just snapshots though. Dolt actually stores the tables, rows, and commits in a Merkle DAG, like Git. So you get branch and merge. You can't do branch and merge with snapshots. (You also get the full git / github toolbox: push and pull, fork and clone, rebase, and most other git commands)
- qbasic_forever 6y agoYeah it's a neat idea but I struggle to think of good use-cases for merge, other than toy datasets. If I'm working on a service that's sharding millions of users across dozens of DB instances a merge is going to be incomprehensible to understand and reason about conflicts.
- MyNameIsFred 6y ago> Yeah it's a neat idea but I struggle to think of good use-cases [...] If I'm working on a service [...] I suspect that's simply not the use-case they're targeting. You're thinking of a database as simply the persistence component for your service, a means to an end. For you, the service/app/software is the thing you're trying to deliver. Where this looks useful is the cases where the data itself is the thing you're trying to deliver.
- yarg 6y agoMerging is hard, but the rest can be done with copy-on-write cloning (or am I missing something?).
- laurent92 6y agoWordpress would have benefited from this. What a lot of webmasters want is, test the site locally, then merge it back. A lot of people turned to Jekyll or Hugo for the very reason that it can be checked into git, and git is reliable. A static website can’t get hacked, whereas anyone who has been burnt with Wordpress security fail knows they’d prefer a static site. And even more: People would like to pass the new website from the designer to the customer to managers — Wordpress might have not needed to develop their approval workflows (permission schemes, draft/preview/publish) if they had had a forkable database.
- deleted 6y ago[deleted]
- kenniskrag 6y agothey have a theme preview now. :)
- TimTheTinker 6y agoSure, but Wordpress is still running PHP files for every page load and back-end query. Dolt would help offload some of the code’s complexity, but that would still leave a significant attack surface area. In other words, by itself Dolt couldn’t solve the problem of Wordpress not being run mostly from static assets on a server (plus an API).
- berkes 6y agoAt the root of this lies the problem that content, configuration and code is stored in one blurp (a single db). Never clearly bounded nor contained. Config is spread over tables. Usergenerated content (from comments to orders) mixed up with redactional content: often even in tables, sometimes even in the same column. Drupal suffers the same. What is needed, is a clear line between configuration (and logic), redactional content, and user generated content. Those three have very distinct lifecycles. As long as they are treated as one, the distinct lifecycles will conflict. No matter how good your database merging is: the root is a missing piece of architecture.
- kyrieeschaton 6y agoPeople interested in this approach should compare Rich Hickey's Datomic.
- einpoklum 6y agoBut do you really need this functionality, if you already have an SQL database? That is, you can: 1. Create a table with an extra changeset id column and a branch id column, so that you can keep historical values. 2. Have a view on that table with the latest version of each record on the master branch. 3. Express branching-related actions as actions on the main table with different record versions and branch names 4. For the chocolate sprinkles, have tables with changeset info and branch info and that gives you a poor man's git already - doesn't it?
- resonantjacket5 6y agoThat doesn't let you merge in someone else's changes easily. Aka two team members make changes to the (local copy of) database and now you want to merge it. I mean sure you can have another database tracking their changes and a merging algorithm but that's what dolt is doing for you
- joshspankit 6y agoI never understood why we don’t have SQL databases that track all changes in a “third dimension” (column being one dimension, row being the second dimension). It might be a bit slower to write, but hook the logic in to write/delete, and suddenly you can see exactly when a field was changed to break everything. The right middleware and you could see the user, IP, and query that changed it (along with any other queries before or after).
- tthun 6y agoMS SQL server 2016 onwards has temporal tables that support this (point in time data)
- joshspankit 6y agoHuh. I just read the spec. Not quite three 'dimension', but looks like exactly what I was asking for: a (reasonably) automatic and transparent record of previous values, as well as timestamps for when they changed. I'll call this a "you learn something every day" and a "hey thanks @tthun (and @predakanga)"
- predakanga 6y agoThis does exist, though support for it is pretty sparse; it's called "Temporal Tables" in the SQL:2011 standard - https://sigmodrecord.org/publications/sigmodRecord/1209/pdfs/07.industry.kulkarni.pdf https://sigmodrecord.org/publications/sigmodRecord/1209/pdfs... Last time I checked, it was supported in SQL server and MariaDB, and Postgres via an extension.
- kenniskrag 6y agoBecause you can do that with after update triggers or server-side in software.
- mcrutcher 6y agoThis has existed for a very long time as a data modeling strategy (most commonly, a "type 2 dimension") and is the way that all MVCC databases work under the covers. You don't need a special database to do this, just add another column to your database and populate it with a trigger or on update.
- Ericson2314 6y agoWhat people usually miss about these things is normal version control benefits hugely from content addressing and normal forms. The salient aspect of relational data is that it's cyclic, this makes content addressing unable to provide normal forms on it's own (unless someone figures out how to Merkle cylic graphs!), but the normal form can still made other ways. The first part is easier enough, store rows in some order. The second part is more interesting: making the choice of surrogate keys not matter (quotienting it away). Sorting table rows containing surrogate keys depending on the sorting of table rows makes for some interesting bags of constraints, for which there may be more than one fixed point. Example: CREATE TABLE Foo ( a uuid PRIMARY KEY, b text, best_friend uuid REFERENCES Foo(b) ); DB 0: 0 Alice 0 1 reclusive Alice, best friends with herself. Just fine. 0 Alice 1 1 Alice 1 2 reclusive Alices, both best friends with the second one. The alices are the same up to primary keys, but while primary keys are to be quotiented out, primary key equality isn't, so this is valid. And we have an asymmetry by which to sort. 0 Alice 1 1 Alice 0 2 reclusive Alices, each best friends with the other. The Alices are completely isomorphic, and one notion of normal forms would say this is exactly the same as DB 0: as if this is reclusive Alice in a fun house of mirrors. All this is resolvable, but it's subtle. And there's no avoiding complexity. E.g. if one wants to cross reference two human data entries who each assigned their own surrogate IDs, this type of analysis must be done. Likewise when merging forks of a database. I'd love be wrong, but I don't think any of the folks doing "git for data" are planning their technology with this level of mathematical rigor.
- TOGoS 6y agoA lot of your comment went over my head, but I have modeled relational data in a way conducive to being stored in a Merkle tree. The trick being that every entity in the system ended up having two IDs. A hash ID, identifying this specific version of the object, and an entity ID (probably a UUID or OID), which remained constant as new versions were added. In a situation where people can have friends that are also people, they are friends with the person long-term, not just a specific version of the friend, so in that case you'd use the entity IDs. Though you might also include a reference to the specific version of the person at the point in time at which they became friends, in which case they would necessarily reference an older version of that person. If you friend yourself, you're actually friending that person a moment ago. A list of all current entities, by hash, is stored higher up the tree. Whether it's better that objects themselves store their entity ID or if that's a separate data structure mapping entity to hash IDs depends on the situation. On second reading I guess your comment was actually about how to come up with content-based IDs for objects. I guess my point was that in the real world you don't usually need to do that, because if object identity besides its content is important you can just give it an arbitrary ID. How often does the problem of differentiating between a graph containing identical Alices vs one with a single self-friending Alice actually come up? Is there any way around it other than numbering the Alices?
- qwerty456127 6y agoI would rather have a "GitHub for data" - an SQL database I could have hosted for free, give everybody read-only access to and give some people I choose R/W access to. That's a thing I miss really.
- Aeolun 6y agoWhat does this look this look like when merging? Does it need a specialized tool? I’m a bit sad it doesn’t seem to have git style syncing though.
- zachmu 6y agoMerge is built in, same as with git. Same syntax too: dolt checkout -b <branch> dolt merge <branch> What do you mean git-style syncing? It has `push`, `pull`, and `fetch`.
- Aeolun 6y agoIn regards to merging, I understand that part, but say I have a conflict, how does it get resolved. Does it open some sort of text editor where I delete the unwanted lines? I mean, if I have a repository on one PC, I can clone from there on a different PC. As far as I can see this only allows S3/GC or Dolthub remotes.
- zachmu 6y agoMerge conflicts get put into a special conflicts table. You have to resolve them one by one just like text conflicts from a git merge. You can do this with the command line or with SQL. It's true we don't support SSH remotes yet. Wouldn't be too hard to add though. File an issue if this is preventing you from using the product. Or use sshfs, which is a pretty simple workaround.
- herpderperator 6y agoThis is pretty cool. I wish `dolt diff` would use + and - though (isn't that standard?) rather than > and < which is harder to distinguish.
- guerrilla 6y agoBoth are "standard" from the UNIX/GNU diff(1) tool. The default behavior gives you the '>' and '<' format and using -u gives you the "unified" '+' and '-' format.
- pgt 6y agoDatomic: https://www.youtube.com/watch?v=Cym4TZwTCNU https://www.youtube.com/watch?v=Cym4TZwTCNU
- deleted 6y ago[deleted]
- scottmcdot 6y agoDolt might be good but never underestimate the power of Type 2 Slowly Changing Dimension tables [1]. For example, if you had an SSIS package that took CSV and imported them into a database, and one day you noticed it accidently rounded the value incorrectly, you could fix the data and retain traceability of the data which was there originally. E.g., SSIS package writes row of data: https://imgur.com/DClXAi5 https://imgur.com/DClXAi5 Then a few months later (on 2020-08-15) we identify that trans_value was imported incorrectly so we update it: https://imgur.com/wdQJWm4 https://imgur.com/wdQJWm4 Then whenever we SELECT from the table we always ensure we are extracting "today's" version of the data: select * from table where TODAY between effective_from and effective_to [1] https://en.wikipedia.org/wiki/Slowly_changing_dimension https://en.wikipedia.org/wiki/Slowly_changing_dimension
- sixdimensional 6y agoI definitely agree, just tossing in the superset concept that Dolt and Type 2 SCD involve - temporal databases [1]. I think the idea of a "diff" applied to datasets is quite awesome, but even then, we kind of do that with databases today with data comparison tools - it's just most of them are not time aware, rather they are used to compare data between two instances of the data in different databases, not at two points in time in the same database. [1] https://en.wikipedia.org/wiki/Temporal_database https://en.wikipedia.org/wiki/Temporal_database
- zachmu 6y agoIf all you want is AS OF semantics, then SCD2 is a great match. Used it a ton in application development myself. Dolt actually makes branch and merge possible, totally different beast.
- skybrian 6y agoThe commit log in Dolt is edit history. (When did someone import or edit the data? Who made the change?) It's not about when things happened. To keep track of when things happened, you would still need date columns to handle it. But at least you don't need to handle two-dimensional history for auditing purposes. So, in your example, I think the "effective" date columns wouldn't be needed. They have ways to query how the dataset appeared at some time in the past. However, with both data changes and schema changes being mixed together in a single commit log, I could see this being troublesome. I suppose writing code to convert old edit history to the new schema would still be possible, similar to how git allows you to create a new branch by rewriting an existing one.
- xkvhs 6y agoWell, maybe this is the ONE. But I've heard "git for data" way too many times to jump on board of the latest tool. I'll wait till "dust settles" and there's a clear winner. Till then, it's parquet, or even pickle.
- crazygringo 6y agoThis is absolutely fascinating, conceptually. However, I'm struggling to figure out a real-world use case for this. I'd love if anyone here can enlighten me. I don't see how it can be for production databases involving lots of users, because while it seems appealing as a way to upgrade and then roll back, you'd lose all the new data inserted in the meantime. When you roll back, you generally want to roll back changes to the schema (e.g. delete the added column) but not remove all the rows that were inserted/deleted/updated in the meantime. So does it handle use cases that are more like SQLite? E.g. where application preferences, or even a saved file, winds up containing its entire history, so you can rewind? Although that's really more of a temporal database -- you don't need git operations like branching. And you really just need to track row-level changes, not table schema modifications etc. The git model seems like way overkill. Git is built for the use case of lots of different people working on different parts of a codebase and then integrating their changes, and saving the history of it. But I'm not sure I've ever come across a use case for lots of different people working on the data and schema in different parts of a database and then integrating their data and schema changes. In any kind of shared-dataset scenario I've seen, the schema is tightly locked down, and there's strict business logic around who can update what and how -- otherwise it would be chaos. So I feel like I'm missing something. What is this actually intended for? I wish the site explained why they built it -- if it was just "because we can" or if projects or teams actually had the need for git for data?
- jorgemf 6y agoMachine Learning. I don't think it has many more use cases
- sixdimensional 6y agoOr more simply put, how about table-driven logic in general? It doesn't have to be as complex as machine learning. There are more use cases than just machine learning, IMHO.
- jedberg 6y agoSuch as? I'm having difficulty coming up with any myself.
- zomglings 6y agoIf anyone from the dolt team is reading this, I'd like to make an enquiry: At bugout.dev, we have an ongoing crawl of public GitHub. We just created a dataset of code snippets crawled from popular GitHub repositories, listed by language, license, github repo, and commit hash and are looking to release it publicly and keep it up-to-date with our GitHub crawl. The dataset for a single crawl comes in at about 60GB. We uploaded the data to Kaggle because we thought it would be a good place for people to work with the data. Unfortunately, the Kaggle notebook experience is not tailored to such large datasets. Our dataset is in a SQLite database. It takes a long time for the dataset to load into Kaggle notebooks, and I don't think they are provisioned with SSDs as queries take a long time. Our best workaround to this is to partition into 3 datasets on Kaggle - train, eval, and development, but it will be a pain to manage this for every update, especially as we enrich the dataset with results of static analysis, etc. I'd like to explore hosting the public dataset on Dolthub. If this sounds interesting to you please, reach out to me - email is in my HN profile.
- zomglings 6y agoThis is the dataset on Kaggle - https://www.kaggle.com/simiotic/github-code-snippets https://www.kaggle.com/simiotic/github-code-snippets
- justinclift 6y agoYeah, that sized database is likely to be a challenge unless the computer system it's running on has scads of memory. One of my projects (DBHub.io) is putting effort towards working through the problem with larger sized SQLite databases (~10GB), and that's mainly through using bare metal hosts with lots of memory. eg 64GB, 128GB, etc. Putting the same data into PostgreSQL, or even MySQL, would likely be much more efficient memory wise. :)
- zomglings 6y agoCan't beat SQLite for distribution as a public dataset, though.
- dang 6y agoSome related past threads: Dolt is Git for data - https://news.ycombinator.com/item?id=22731928 https://news.ycombinator.com/item?id=22731928 - March 2020 (191 comments) Git for Data – A TerminusDB Technical Paper [pdf] - https://news.ycombinator.com/item?id=22045801 https://news.ycombinator.com/item?id=22045801 - Jan 2020 (5 comments) Ask HN: Would you use a “git for data”? - https://news.ycombinator.com/item?id=11537934 https://news.ycombinator.com/item?id=11537934 - April 2016 (10 comments)
- LukeEF 6y agothanks for the mention of TerminusDB (https://github.com/terminusdb/terminusdb https://github.com/terminusdb/terminusdb) -> we are the graph cousins of Dolt. Great to see so much energy in the version control database world! We are currently focused on data mesh use cases. Rather than trying to be GitHub for Data, we're trying to be YOUR GitHub for Data. Get all that good git lineage, pull, push, clone etc. and have data producers in your org own their data. We see lots of organizations with big 'shadow data' problems and data being centrally managed rather than curated by domain experts.
- weeboid 6y agoScanned for this comment before making it. This optically reads as "Dolt", not "DoIt"
- tjkrusinski 6y agoAny info on how this differs from Dat?
- mayureshkathe 6y agoSince you've hosted it on GitHub, would you also consider providing a JavaScript client library for read access? Could pave the way for single page web applications (hosted via GitHub Pages) working with Dolt as their database of choice.
- jonnycomputer 6y agoNaive question here. Aside from it being mysql, what is different here than just using git + sqlite. Update: When I posted, I'd forgotten that SQLite db file is a binary. Not sure what I was thinking.
- bargle0 6y agoMerging a SQLite database is challenging.
- justinclift 6y agoWe do it on DBHub.io. :) If you're into Go, these are the commits where the pieces were hooked together. * https://github.com/sqlitebrowser/dbhub.io/commit/6c40e051ff7b27c6c66c3c7bf5f75c5e2990d17c https://github.com/sqlitebrowser/dbhub.io/commit/6c40e051ff7... * https://github.com/sqlitebrowser/dbhub.io/commit/989ee0d08e6b6888edc5579d3c43bfc362ec9782 https://github.com/sqlitebrowser/dbhub.io/commit/989ee0d08e6...
- unnouinceput 6y agoI'd venture more and say merging any DB is challenging, SQLite or not.
- deleted 6y ago[deleted]
- Nican 6y agoI love the idea of this project so much. Being able to more easily share comprehensible data is an interest of mine. It is not the first time I have seen immutable B-trees being used as a method for being able to query a dataset on a different point in time. Spanner (and its derivatives) uses a similar technique to ensure backup consistency. Solutions such as CockroachDB also allows you to query data in the past [1], and then uses a garbage collector to delete older unused data. The Time-to-live of history data is configurable. [1] https://www.cockroachlabs.com/docs/stable/as-of-system-time.html https://www.cockroachlabs.com/docs/stable/as-of-system-time....
- Nican 6y agoAlbeit, now it makes me wonder how much of the functionality of dolt is possible to be replicated with CockroachDB. The internal data structures of both databases are mostly similar. You can have infinite point-time-querying by setting the TTL of data to forever. You have the ability to do distributed BACKUP/IMPORT of data (Which is mostly a file copy, and also includes historical data) A transaction would be the equivalent of a commit, but I do not think there is a way to list out all transactions on CRDB, so that would have to be done separately. And gain other benefits, such as distributed querying, and high availability. I just find interesting that both databases (CockroachDB and Dolt) share the same principal of immutable B-Trees.
- smarx007 6y agoIf you are interested to one step beyond, you can check out https://terminusdb.com/docs/terminusdb/#/ https://terminusdb.com/docs/terminusdb/#/ which uses RDF to represent (knowledge) graphs (and has revision control).
- 1f60c 6y agoI'm not sure if this is supposed to be Dolt or DoIt, but using a swear word for a name (even a relatively mild one) is pretty distracting, IMHO.
- drewwwwww 6y agopresumably a riff on git, the well known famously unsuccessful version control system
- 1f60c 6y agoHuh, I had no idea. (I'm not a native speaker.)
- stjohnswarts 6y agoDoIt isn't really a swear word. It's a euphemism for coitus. It's also a very common phrase in general so it really isn't strongly associated with "doing it" to cause any issues with English speakers. Aka it's fine as a project name. See Nike's "Just do it." advertisement campaign, they would have never gone with the phrase if it had strong negative connotations.
- zachmu 6y agoIt's DOLT. San serif fonts.
- efikoman 6y ago?
- devwastaken 6y agoSomething like this but for sqlite would be great for building small git enabled applications/servers that cana benefit from the features git provides but only need a database and a library to do it.
- lifty 6y agoNoms might be what you’re looking for (https://github.com/attic-labs/noms https://github.com/attic-labs/noms). Dolt is actually a fork of Noms.
- deleted 6y ago[deleted]
- px43 6y agoIs there a nice way to distribute a database across many systems? Would this be useful for a data-set like Wikipedia? It would be really nice if anyone could easily follow Wikipedia in such a way that anyone could fork it, and merge some changes, but not others, etc. This is something that really needs to happen for large, distributed data-sets. I really want to see large corpuses of data that can have a ton of contributors sending pull requests to build out some great collaborative bastion of knowledge.
- metaodi 6y agoThis is basically Wikidata [1], where the distribution is linked data/RDF with a triple store and federation to query different triple stores at once using SPARQL. [1] https://wikidata.org https://wikidata.org
- 0xbadcafebee 6y agoIf this could somehow work with existing production MariaDB/PostgreSQL databases, this would be the next Docker. Sad that it requires its own proprietary DB (SQL compatible though it may be, it's not my existing database). I wish something completely free like this existed just to manage database migrations. I don't think anything is quite this powerful/user-friendly (if you can call Git that) and free.
- jka 6y agoAre you familiar with SQLAlchemy Alembic[1]? For best results it requires modeling your project in Python using SQLAlchemy, and even then, some migrations may require manual tweaks. But it is free and supports an extensive range of database systems. [1] - https://alembic.sqlalchemy.org/en/latest/ https://alembic.sqlalchemy.org/en/latest/
- anasbarg 6y agoI like this. A while ago, I was asking my co-founder the question "Why isn't there Git for data like there is for code?" while working on a database migrations engine that aims to provide automatic database migrations (data & schema migrations). After all, code is data.
- controlledchaos 6y agoThis is exactly what I was looking for!
- rymurr 6y agoI've been working on https://github.com/projectnessie/nessie https://github.com/projectnessie/nessie for about a year now. Its similar to Dolt in spirit but aimed at big data/data lakes. Would welcome feedback from the community. Its very exciting to see this field picking up speed. Tons of interesting problems to be solved :-)
- knbknb 6y agoIs something like "dolt diff master..somebranch -- mytable" possible? Same question: Are in-between branch comparisons possible in dolt? Which "diff" subcommands to you plan to support in the future?
- mleonhard 6y agoThis could be useful for integration tests, especially for testing schema changes.
- valtism 6y agoI've looked into Dolt perviously for powering our product that is Git for financials. The trouble is that most of our complexity lies in our data's relationships. Merging is very tricky when some data that has been added in one branch has not had a specific change in properties when other data in a master branch has been modified.
- LukeEF 6y agoGood point - and why a graph based approach has advantages in some use cases (disclaimer TerminusDB co-founder here). Allows more flexibility in versioning relationships and more scope to merge in the scenario you outline.