3 ms·
Here is one of the DVC maintainers :) Congrats! It's great to see more tools for codifying data in different scenarios. To be honest, since you introduce a new
by ishcheklein 6y ago
Here is one of the DVC maintainers :) Congrats! It's great to see more tools for codifying data in different scenarios.
To be honest, since you introduce a new workflow and a few new concepts it's not that easy to get the right perspective in 5 minutes (I know the same problems exists with DVC and we've been iterating on docs a lot). Mind a few questions?
Do I understand it right, that is mostly focused on tabular data? Kinda git checkout for an SQL table?
- chatmasta 6y agoHey! Thanks for the kind words. Yeah, it's a small space and we all seem to be aware of each other. In fact, I don't know if you saw it, but we do briefly mention dvc in the FAQ [0]. > Kinda git checkout for an SQL table This is basically right, yes. We implement some Git- and Docker- like operations on top of the SQL standard. For example, we borrow the idea of "delta compression" (storing only changes) from Git. Like Git, we store commits as a set of objects representing the changes since the last commit. But whereas Git versions files with lines as the unit of change, Splitgraph versions tables with rows as the unit of change. Our "objects" are actually cstore files that represent fragments of a table, with content addressable hashes generated with LTHash [1]. If you want to read about this in more detail, have a look at the documentation for "objects." [2] [0] https://www.splitgraph.com/docs/getting-started/frequently-asked-questions#dvc-datalad- https://www.splitgraph.com/docs/getting-started/frequently-a... [1] LTHash has the useful property that the sum of all hashes of individual fragments composing a table is equal to the content hash of a whole table. This is how we are able to have content addressable objects, which unlocks a lot of tricks in terms of efficiency and allows techniques like "layered querying". [2] https://www.splitgraph.com/docs/concepts/objects https://www.splitgraph.com/docs/concepts/objects
- ishcheklein 6y agoThanks! Still trying to wrap my mind around use cases :) What is the main adoption path for the tool do you see? People who use flat files and they are large enough and/or need some provenance? Or people who already have Postgres and they want to version it? Or is it about production use case (like Docker) - make snapshot to deliver it consistently? Btw, regarding the FQA - thanks! From my take on Splitgraph though one the main DVC's difference is that it deals with flat files. From software engineering perspective - it is very close to Git lfs on steroids + some higher level features similar to makefiles, ML metrics, etc.
- chatmasta 6y ago> What is the main adoption path for the tool do you see? We hope that anyone who works with data on a daily basis can benefit from Splitgraph. When we first started, we gave a presentation at a Docker meetup called "Docker for Data" [0] with the idea that we wanted to do for data scientists what Docker did for DevOps. We hope that people will be able to throw out some of their fragile ETL scripts in favor of using Splitfiles, just like DevOps engineers could throw out their Salt and Chef scripts when they switched to Docker. > Or people who already have Postgres and they want to version it? It's important to note that in many cases, Postgres is really just an intermediary. Splitgraph allows you to "mount" data from any source, not just Postgres databases, by leveraging Postgres foreign data wrappers (FDWs) [1]. This idea of "mounting" is one of the core abstractions of Splitgraph that makes it really powerful, because it lets you use a common format (Splitfiles) to transform and query data from anywhere. And once you've built an image from a bunch of disparate data sources, you can use a core set of your favorite tools to query it, since as far as they're concerned, it's just a Postgres schema. So in this sense, Splitgraph can serve as a sort of universal translation layer for data from all over the place. For a really powerful example using this idea, see the example where we mount two tables from two separate data portals (Chicago and Cambridge), and join between them. [2] [0] We gave two presentations in 2018, both similar. A lot of details have changed since then, but the core abstractions are the same, and they might give some insight into our direction: [0.a] https://www.slideshare.net/splitgraph/splitgraph-docker-for-data-119112722 https://www.slideshare.net/splitgraph/splitgraph-docker-for-... [0.b] https://www.slideshare.net/splitgraph/splitgraph-ahl-talk https://www.slideshare.net/splitgraph/splitgraph-ahl-talk [1] https://www.splitgraph.com/docs/ingesting-data/foreign-data-wrappers/introduction https://www.splitgraph.com/docs/ingesting-data/foreign-data-... [2] https://www.splitgraph.com/docs/ingesting-data/socrata#using-metabase-to-join-and-plot-data-from-multiple-data-portals https://www.splitgraph.com/docs/ingesting-data/socrata#using...
- ishcheklein 6y agoOkay, thanks! I think I got the idea. Have you seen Quilt btw? They had the same message initially - Docker for Data, packaging data, etc. Implementation was very different though. I somewhat don't like this analogy btw, see how you mentioned Docker changed life for DevOps in the first place (vs engineers), the same here - data scientists don't care about packaging data - there should be strong incentive to do so. A few specific questions: 1. where does query execution happen - always client or remote as well? 2. in a global "Github for data" case is there some discovery mechanism for existing data? 3. do you provide a public storage to cover the case for Github for data? or is it now more like torrent - peers host and pay for data storage? Btw, what is major direction for you - Github (public collaboration and sharing) or internal versioned warehouses (or some other internal case?).