4 ms·
Show HN: Splitgraph - Build and share data with Postgres, inspired by Docker/Git
- chatmasta 6y agoHey HN! I’m Miles, co-founder of Splitgraph along with Artjoms (mildbyte). We met back in 2018, when I reached out to him after reading his blog on HN and realizing we lived right next to each other. Neither of us had a “real job,“ and we both wanted to build something truly innovative and cool. We tossed around a few ideas, but ultimately we couldn’t resist the idea of building “GitHub for data,” which seemed like an obvious gap in the market. After nearly two years of development, we are finally ready — and extremely excited — to share it with the world. We are not the first to notice this gap or try to build this product. So we wanted to make sure we did it right. We made sure to start from “first principles” and really analyze the problem space. We ended up realizing that it’s not strictly Git or GitHub that people want “for data.” Rather, people just want to be able to work with data as easily as they can work with code. They want to experiment, build and maintain data without needless overhead. Tools like Git and Docker are ubiquitous in any software engineer’s workflow, and we took a lot of inspiration from them when designing Splitgraph. We thought about why people like and use these tools, and tried to translate their benefits to the domain of data science. Our core philosophy is to stay out of the way, and work with existing abstractions instead of introducing new ones. You can version your code with Git without switching filesystems. You can build Docker images without changing your code to work in Docker. Our goal with Splitgraph is to provide an easy path to incremental adoption, so you can introduce it into your existing workflows where and when it makes sense. Splitgraph is powered by Postgres, and provides an easy way to build and share versioned datasets, along with a whole bunch of other benefits. We encourage you to read the landing page which (hopefully) explains it well. The documentation goes into much more detail, and if you have ten minutes and Docker installed, you can try Splitgraph for yourself. [0] If you work with data, we really hope you’ll give Splitgraph a try. We’re here to answer any questions, and we’ve also created a Discord server [1] to hopefully build a bit of a community around Splitgraph. [0] https://www.splitgraph.com/docs/getting-started/five-minute-demo https://www.splitgraph.com/docs/getting-started/five-minute-... [1] https://discord.gg/eFEFRKm https://discord.gg/eFEFRKm
- philips 6y agoThis is so cool! I have been looking around for databases that have any sort of cryptographic digest of data to ensure integrity. And this is the first time I have seen something do that. Could the snapshots and content addressability be used for regular backups of application databases?
- mildbyte 6y agoThanks, glad you like it! Theoretically, yes, you can use Splitgraph as a PostgreSQL replication client and then occasionally run a commit to produce a new delta. But we're currently focused on the OLAP use case and so Splitgraph images don't yet support storing things like indexes, views, triggers, stored procedures etc.
- philips 6y agoVia their (very good FAQ) Writing to PostgreSQL tables that are change-tracked by Splitgraph is almost 2x slower than writing to untracked tables (Splitgraph uses audit triggers to record changes rather than diffing the table at commit time).
- username3 6y agoHow does this compare to Dolt and DoltHub?
- timsehn 6y agoCEO of the company that built Dolt here. I just did some preliminary research and it seems that we have a similar mission. We both want to version control data. The thing Dolt does that splitgraph does not do is support branches, diffs, merges and conflicts on data and schema. Dolt has its own storage engine to do this efficiently whereas Splitgraph relies on Postgres. Having native Postgres with a versioning layer on top has other advantages so we're excited to see how Splitgraph's approach works in practice. Excited to play with it. We love to see more tools in this underserved space.
- timsehn 6y agoAlso a survey of other "Git for data" options from a few months ago: https://www.dolthub.com/blog/2020-03-06-so-you-want-git-for-data/ https://www.dolthub.com/blog/2020-03-06-so-you-want-git-for-...
- mildbyte 6y agoThere aren't many places where Splitgraph intersects with Dolt. Dolt aims to build a database from the ground up to have Git semantics and a real commit graph, whereas Splitgraph works on top of an existing RDBMS (PostgreSQL) and performs its operations by manipulating database tables. Here's a quick overview of differences where we do intersect. With data versioning, we cherry-picked (no pun intended!) only a few concepts/commands from Git that we think are applicable. In particular, like Tim mentioned, we don't support cell-level merges. However, we offer a higher level DSL (Splitfiles[-1]) to collaborate on data. After a Splitgraph image is checked-out, it's just a set of ordinary tables, so you get PostgreSQL feature parity and read-write performance out of the box. That means we're compatible with any existing PostgreSQL clients (DataGrip, pgcli, pgAdmin, DBeaver...), tools (Metabase, dbt...) and extensions, including those that define custom types (e.g. PostGIS for geospatial data)[0]. We also offer a way of querying Splitgraph images without checking them out[1]. This lets it download required table regions on the fly (IIRC with Dolt, you have to fully clone the whole dataset and all of its history before querying it) and in a lot of cases can be faster than PostgreSQL itself[2]. This is also completely transparent to and compatible with existing clients. Splitgraph is decentralized and can treat any instance as a remote and push data there, with authorization provided by PostgreSQL, so you can use methods like LDAP/RADIUS/Kerberos to control access. It looks like you can spin up a Dolt remote with [3] but it's not very well documented so I can't comment on it. Finally, we offer a lot of features beyond data versioning (reproducible dataset builds with provenance tracking, mounting other databases, automatically generated REST API for all datasets on Splitgraph Cloud etc). When I tried it out a couple of months ago, Dolt's MySQL server didn't work with mysql_fdw. But their MySQL compatibility is progressing at an impressive pace and if it works now, you'll be able to query Dolt datasets directly from Splitgraph (by running sgr mount mysql[4]) and use them in Splitfiles. In the meantime, you can import data from Dolt into Splitgraph by using their dolt-to-pg adapter. [-1] https://www.splitgraph.com/docs/concepts/splitfiles https://www.splitgraph.com/docs/concepts/splitfiles [0] https://www.splitgraph.com/product/splitgraph/integrations https://www.splitgraph.com/product/splitgraph/integrations [1] https://www.splitgraph.com/docs/large-datasets/layered-querying https://www.splitgraph.com/docs/large-datasets/layered-query... [2] https://github.com/splitgraph/splitgraph/blob/master/examples/benchmarking/benchmarking_real_data.ipynb https://github.com/splitgraph/splitgraph/blob/master/example... [3] https://github.com/liquidata-inc/dolt/tree/master/go/utils/remotesrv https://github.com/liquidata-inc/dolt/tree/master/go/utils/r... [4] https://www.splitgraph.com/docs/ingesting-data/foreign-data-wrappers/load-mysql-tables https://www.splitgraph.com/docs/ingesting-data/foreign-data-...
- ishcheklein 6y agoHere is one of the DVC maintainers :) Congrats! It's great to see more tools for codifying data in different scenarios. To be honest, since you introduce a new workflow and a few new concepts it's not that easy to get the right perspective in 5 minutes (I know the same problems exists with DVC and we've been iterating on docs a lot). Mind a few questions? Do I understand it right, that is mostly focused on tabular data? Kinda git checkout for an SQL table?
- chatmasta 6y agoHey! Thanks for the kind words. Yeah, it's a small space and we all seem to be aware of each other. In fact, I don't know if you saw it, but we do briefly mention dvc in the FAQ [0]. > Kinda git checkout for an SQL table This is basically right, yes. We implement some Git- and Docker- like operations on top of the SQL standard. For example, we borrow the idea of "delta compression" (storing only changes) from Git. Like Git, we store commits as a set of objects representing the changes since the last commit. But whereas Git versions files with lines as the unit of change, Splitgraph versions tables with rows as the unit of change. Our "objects" are actually cstore files that represent fragments of a table, with content addressable hashes generated with LTHash [1]. If you want to read about this in more detail, have a look at the documentation for "objects." [2] [0] https://www.splitgraph.com/docs/getting-started/frequently-asked-questions#dvc-datalad- https://www.splitgraph.com/docs/getting-started/frequently-a... [1] LTHash has the useful property that the sum of all hashes of individual fragments composing a table is equal to the content hash of a whole table. This is how we are able to have content addressable objects, which unlocks a lot of tricks in terms of efficiency and allows techniques like "layered querying". [2] https://www.splitgraph.com/docs/concepts/objects https://www.splitgraph.com/docs/concepts/objects
- ishcheklein 6y agoThanks! Still trying to wrap my mind around use cases :) What is the main adoption path for the tool do you see? People who use flat files and they are large enough and/or need some provenance? Or people who already have Postgres and they want to version it? Or is it about production use case (like Docker) - make snapshot to deliver it consistently? Btw, regarding the FQA - thanks! From my take on Splitgraph though one the main DVC's difference is that it deals with flat files. From software engineering perspective - it is very close to Git lfs on steroids + some higher level features similar to makefiles, ML metrics, etc.
- zmmmmm 6y agoI'm probably a bit naive about this but could it make it unnecessary to explicitly create database dumps as backups in scenarios where you need a rollback? ie: could I just tag the database and be guaranteed I would later get back that data if, for example, my upgrade failed and I wanted to restore, simply by checking out the tag?
- mildbyte 6y agoThat does sound very ambitious for now! I discussed below that we're focused on the OLAP use case (manage actual data, not DDL around it), so triggers, indexes and functions that you create won't be stored in the Splitgraph image you'll make (a Splitgraph image is not a full database dump). For things like schema migrations, PostgreSQL itself has transactional DDL: column deletions/additions can be wrapped in a transaction, so you can ROLLBACK if your migration fails. This might be more appropriate for your use case? (Note that you can still add DDL commands to a Splitgraph table after you check it out, since at that point it's just a Postgres table. In theory it would be possible to track DDL changes with some other mechanism, and apply them after loading a version of data)
- ahnick 6y agoPersonally I think I'm more drawn to the dotmesh approach (https://docs.dotmesh.com/concepts/architecture/ https://docs.dotmesh.com/concepts/architecture/), but the one problem data has is as it gets massive it becomes really hard to move it around and I guess that's where trying to layer git like workflows on top of it become intractable. It's like data has it's own gravity and often times it is just easier to bring other things to the data, rather than the other way around. IIRC Bryan Cantrill said something similar about data when Joyent was developing their object storage system Manta (https://www.youtube.com/watch?v=79fvDDPaIoY); https://www.youtube.com/watch?v=79fvDDPaIoY); ergo, perhaps the Splitgraph approach will meet with better success.