4 ms·
This seems great, also in performance, but why would you use this over making plain file copy-based backups, since SQLite is single file based? I'd assume this
by phil294 4y ago
This seems great, also in performance, but why would you use this over making plain file copy-based backups, since SQLite is single file based? I'd assume this solution might be preferable because it saves deltas instead of absolute, but the Readme does not say so. If the main advantage here is on-the-fly selecting from different versions, granted, that is unique, even though I cannot really imagine a production use case for this.
So far, I have been using file copies and GFS backup scheme based solutions like Borg, which also does deltas in chunks, compression, and encryption. Perhaps with Borg one should better even backup SQL dumps instead of db files, I don't know.
- devnull3 4y agoOnly modified pages go into a version data file [1] [1] https://github.com/sudeep9/mojo/blob/main/design.md https://github.com/sudeep9/mojo/blob/main/design.md
- arjvik 4y agoCould you achieve the same by using BTRFS to store the db file?
- devnull3 4y agoYes. But you need to have such a filesystem installed at the first place. Mojo can run on any filesystem which does not support snapshots/versions. This makes it portable not only across different fs but also across OS.
- rahimiali 4y agodon't you get the same benefit if you version controlled the db file with git? with git, each commit saves a diff from the previous one as a blob. the difference is that in git, in addition to the diffs, you also have to create a working copy of the db, which means you use up at least 2x the storage your system uses. in your implementation, the diff blobs are the live db, which saves you ~2x storage. is that the main benefit?
- devnull3 4y agoIn git, each version of the database will be a full copy. The git has to perform diff i.e. scan the database file. Imagine doing commits & creating snapshots very frequently. Have a look at https://github.com/sudeep9/mojo/blob/main/design.md#index https://github.com/sudeep9/mojo/blob/main/design.md#index
- rahimiali 4y agosorry, yes, you mention in another comment the use case of multiple readers operating on different versions of the db simultaneously. that'd be difficult to do with git for the reason you mention.
- Lex-2008 4y agore: backup SQL dumps instead of db files - indeed! In my small experiment (SQL databases in a browser profile, [1]), simple `sqlite3 "$file" .dump | gzip >"$file.sql.gz"` decreased size of files to be backed up about 10 times! I'm not sure how it compares to Borg's delta- and zstd-compression, though. [1]: http://alexey.shpakovsky.ru/en/minimizing-size-of-browser-profiles-backups.html http://alexey.shpakovsky.ru/en/minimizing-size-of-browser-pr...
- phil294 4y agohm but can't you do the same thing (gzipping) with the db itself? Edit. Regardless of if you maybe even meant that, I made a small test too: A 42 MB db zipped is 8.8M, but its sql dump zipped is 4.9M, so almost half the size. Pretty good already. If you don't just `.dump` but actually output the tables in a consistently sorted manner, the size might go down even more and will enable very efficient delta-ing. Questionable value to effort ratio though...
- Lex-2008 4y agore: gzipping the db itself - yep, I tried and saw the same effect as you did: about half a size. I think it's because of indexes (I once played with sqlite3_analyzer[1] and on some databases it showed about half of disk space was used by indexes) and unused pages (all the space which can be vacuum'ed - again, size of some databases can be decreased almost twice by vacuum'ing). [1]: https://www.sqlite.org/sqlanalyze.html https://www.sqlite.org/sqlanalyze.html On the other side, indeed, text dump is likely not the most space-efficient way of storing raw DB data (compared to some binary one), so I wouldn't be surprised to find databases for which gzipped sql dump is bigger than (gzipped) database itself. I would even say that I'm surprised that it's not true for most databases :)