4 ms·
One of the benefits of building on Git is a lot of people have put a lot of time into make it work really well with lots of objects. And even though we say "fil
by polemic 3y ago
One of the benefits of building on Git is a lot of people have put a lot of time into make it work really well with lots of objects. And even though we say "files", Git abstracts that into packfiles etc very efficiently.
So, we're seening pretty good performance. We're maintaining a number of repositories with several millions features, with a decade of weekly updates of ~10,000+ rows. It _does_ take some time to push that data around, but it's _vastly_ better than old ways, and once you have your clone, maintaining updates becomes extremely trivial - a _major_ unsolved problem in the GIS/data world.
I'd add - Kart has GIS specific features that nullify some of these issues. The ability to spatially index the objects, then filtering them on Clone, means I rapidly clone a tiny subset of the data to work with.
- everybodyknows 3y ago> The ability to spatially index the objects, then filtering them on Clone, means I rapidly clone a tiny subset of the data to work with. Okay -- so is the "--depth=N" filtering option to git-clone supported as well? And does it remain useful in the context of Kart applications?
- polemic 3y agoYes, you can do shallow clones with `--depth` as well. This is incredibly useful - it means we can publish massive Kart repositories of spatial data with lots of versioning info, but still allow users to work with small subsets of the most recent changes. Very important for typical GIS use cases.
- Mertax 3y agoIs there a public git repo available somewhere that represents a Kart repository? Are the raw files in the working repository GeoPackages? How is it tracking the changes made inside the geopackages? What happens if it's replaced with an updated copy of the geopackage the was edited via some other application? How does it diff the changes?
- amatix 3y agoGood questions > Are the raw files in the working repository GeoPackages? The working copy for a vector/table dataset can be in a GeoPackage or a SQL database like PostGIS. For rasters/point-clouds they're flat files. > How is it tracking the changes made inside the geopackages? In general, triggers which store RowIDs/PKs of inserts/updates/deletes. Then when you ask for a diff or make a commit Kart figures out any actual row-level (or schema) differences. > What happens if it's replaced with an updated copy of the geopackage the was edited via some other application? If it's edited by something else (QGIS, ArcGIS, python/go/whatever application, SQL CLI, whatever) it'll work: you do edits where you want to. If it's replaced by something else, it won't work. > How does it diff the changes? Comparing the features/rows in the repository (and their schemas) against the rows in the working copy database. It uses the stored list of modified rowids to make this fast.