3 ms·
I wish people would explain why they created their project and what their pain points were with existing alternatives. All I can find is "Contrary to existing t
by choward 5y ago
I wish people would explain why they created their project and what their pain points were with existing alternatives. All I can find is "Contrary to existing tools, Atlas intelligently plans schema migrations for you, based on your desired state."
What's so bad about writing an explicit migrations using something like Flyway? I'm a fan of declarative configuration for the most part but it doesn't seem that beneficial here to me. Sure, using it to modify the schema to the desired state is fine from just the schema perspective. But there is also data you have to worry about. Like someone else stated, the rename example is probably the most trivial example. How would the tool know if I wanted to delete one column and add another or if I wanted to rename the column?
Another more complicated example is what if I had a column called created_at that stored a string and I wanted it to store a date instead. How would atlas know that just by changing the type? I might want it deleted and recreated. But if not, how would it know what format I stored the string in and how to parse it?
For the rename problem, each column could have it's own id so if the id doesn't change it's a rename. Or it could ask while running the migrate command and then store those answers somewhere to be used when deploying to production which seems like a pain. Something like that might work for changing the data type too. You have two columns with the same name and different ids. The new column could have something to specify to transform the data from the old column.
(I know it's a terrible idea to change the data type of a column in production, but the same idea applies if you want to copy the data converted to a new type to a new column, get your production code using that new column, then delete the other column.)
Either way, this doesn't seem to by much over the classic migration approach. You still have to think about how the data is going to move or be manipulated. And with the classic approach they generate a schema so you can see the final state of your database and make sure it's what you want. With the atlas approach you have to approved the planned migration which is basically the same thing as making sure the schema generated is correct using the classic imperative approach.
- evanelias 5y agoDeclarative schema management provides some nice properties that aren't present in imperative/migration-based tools. I'm the creator of a widely-used declarative tool for MySQL/MariaDB called Skeema, and I wrote a blog post summarizing some advantages a few years back: https://www.skeema.io/blog/2019/01/18/declarative/ https://www.skeema.io/blog/2019/01/18/declarative/ You are correct that renames are problematic with declarative tools, but in production renames are problematic in general because of deploy-order concerns. Best practice with schema changes is always for applications to be able to work fine with both the old and new schema, and renames typically break this, as most ORMs / database interface layers don't support this. So renames already require a special process at companies that even allow renames in production (in my experience, many do not). Skeema treats them as an out-of-band activity; it gets out of your way and you handle the rename outside of the tool, and then can use Skeema to update your repo afterwards. If you accidentally try to do a rename inside the tool, it will treat it as a drop-and-create, but it prevents the push from proceeding since it detects it is destructive.
- satyrnein 5y agoI find Skeema pretty compelling, but I couldn't figure out one aspect: Say I'm adding a gender column to my employees table. I see how Skeema would be an alternative to running DDL in a migration. But I would end up still needing a migration to backfill the data. So I could use Skeema, but I still need migrations (we also occasionally fix data due to bugs, etc). At that point, we were less motivated to add an additional process. Is that what Skeema users do, or is there some other approach, or maybe this isn't what I should be doing in the first place?
- evanelias 5y agoCurrently, Skeema doesn't interact with DML or provide anything to help here. But that means you're free to use whatever solution you'd like, and Skeema won't get in your way, even if you store your DML in the same repo or even same subdirectories. Part of the reason for this is that some of Skeema's users are quite large, and larger MySQL users already have in-house solutions for row data migrations. When your tables are huge and/or sharded, a data migration consists of a lot more complexity than just putting an UPDATE statement into a .sql file :) I do agree it would be good to have some options in Skeema for DML at various scales, and it's something I plan to start approaching in the future, hopefully later this year. Overall my approach with Skeema's roadmap has been to first get DDL right for tables, then get DDL right for other types of objects (finally complete), and only then move on to considering automation for other areas (whether that be DML, or something like managing users/grants, database global variables, etc).
- diegoveralli 5y agoAt my previous company there were several attempts to build declarative migration systems, one of which was more or less successful. They were all a lot more modest in scope and goals than Atlas. The motivation is very straight forward: if you work in data management with different schemas, you often see similar migration patterns, so it makes sense to try to build abstractions representing these patterns. If you couple this with the usual problem of the migration / schema duality, and some people who want to have a replica of the current state of the schema in the codebase in some way ("desired state"), you can see how we end up with schema-diff migration tools. Whether these tools work well or not I think depends on the rate of schema changes vs size of the data store. The larger the dataset, the more specific you need to be with the migration strategies, so if you're not making schema changes very often, using a "clever", schema-diff migration tool adds risk for little benefit, and you might need to bypass the tool on a regular basis. I don't work in a scenario in which such a tool would bring benefits, but there are projects that have frequent schema changes and data sizes small enough that you don't need to do anything special to maintain availability during migration.
- nijave 5y agoI used to work on a monolithic Rails app and migrations would build up over time. We'd "collapse" migrations by doing a schema dump every couple weeks. Without it, they'd accumulate thousands of migrations which would take quite a while locally and in CI to rebuild the database It's also much easier to programmatically inspect. In ActiveRecord, the model attributes are dynamically determined from the DB (for better or worse) so you need to also take any migrations into account to know what fields a model has. From an inspection standpoint, it's nice to have the desired representation in code