4 ms·
Thanks! Still trying to wrap my mind around use cases :) What is the main adoption path for the tool do you see? People who use flat files and they are large e
by ishcheklein 6y ago
Thanks! Still trying to wrap my mind around use cases :) What is the main adoption path for the tool do you see?
People who use flat files and they are large enough and/or need some provenance?
Or people who already have Postgres and they want to version it?
Or is it about production use case (like Docker) - make snapshot to deliver it consistently?
Btw, regarding the FQA - thanks! From my take on Splitgraph though one the main DVC's difference is that it deals with flat files. From software engineering perspective - it is very close to Git lfs on steroids + some higher level features similar to makefiles, ML metrics, etc.
- chatmasta 6y ago> What is the main adoption path for the tool do you see? We hope that anyone who works with data on a daily basis can benefit from Splitgraph. When we first started, we gave a presentation at a Docker meetup called "Docker for Data" [0] with the idea that we wanted to do for data scientists what Docker did for DevOps. We hope that people will be able to throw out some of their fragile ETL scripts in favor of using Splitfiles, just like DevOps engineers could throw out their Salt and Chef scripts when they switched to Docker. > Or people who already have Postgres and they want to version it? It's important to note that in many cases, Postgres is really just an intermediary. Splitgraph allows you to "mount" data from any source, not just Postgres databases, by leveraging Postgres foreign data wrappers (FDWs) [1]. This idea of "mounting" is one of the core abstractions of Splitgraph that makes it really powerful, because it lets you use a common format (Splitfiles) to transform and query data from anywhere. And once you've built an image from a bunch of disparate data sources, you can use a core set of your favorite tools to query it, since as far as they're concerned, it's just a Postgres schema. So in this sense, Splitgraph can serve as a sort of universal translation layer for data from all over the place. For a really powerful example using this idea, see the example where we mount two tables from two separate data portals (Chicago and Cambridge), and join between them. [2] [0] We gave two presentations in 2018, both similar. A lot of details have changed since then, but the core abstractions are the same, and they might give some insight into our direction: [0.a] https://www.slideshare.net/splitgraph/splitgraph-docker-for-data-119112722 https://www.slideshare.net/splitgraph/splitgraph-docker-for-... [0.b] https://www.slideshare.net/splitgraph/splitgraph-ahl-talk https://www.slideshare.net/splitgraph/splitgraph-ahl-talk [1] https://www.splitgraph.com/docs/ingesting-data/foreign-data-wrappers/introduction https://www.splitgraph.com/docs/ingesting-data/foreign-data-... [2] https://www.splitgraph.com/docs/ingesting-data/socrata#using-metabase-to-join-and-plot-data-from-multiple-data-portals https://www.splitgraph.com/docs/ingesting-data/socrata#using...
- ishcheklein 6y agoOkay, thanks! I think I got the idea. Have you seen Quilt btw? They had the same message initially - Docker for Data, packaging data, etc. Implementation was very different though. I somewhat don't like this analogy btw, see how you mentioned Docker changed life for DevOps in the first place (vs engineers), the same here - data scientists don't care about packaging data - there should be strong incentive to do so. A few specific questions: 1. where does query execution happen - always client or remote as well? 2. in a global "Github for data" case is there some discovery mechanism for existing data? 3. do you provide a public storage to cover the case for Github for data? or is it now more like torrent - peers host and pay for data storage? Btw, what is major direction for you - Github (public collaboration and sharing) or internal versioned warehouses (or some other internal case?).
- chatmasta 6y agoExcellent questions, thank you. > Have you seen Quilt btw? Yes, we have. It seems in this space that everything has been done or pitched before, but in our opinion nobody has hit the exact right execution yet. A big problem with a lot of existing tools is that they disrupt your workflow, or are otherwise hard to adopt without major adjacent changes. Our core philosophy with Splitgraph is to stay out of the way. As long as we can keep this up, and as long as we can continue building on a core set of simple abstractions, we think we stand a pretty good chance. > data scientists don't care about packaging data Indeed. It's worth noting that packaging data with Splitfiles is entirely optional. You can also run ad-hoc queries against a database with change-tracking enabled (meaning, Splitgraph audit triggers are installed), and periodically commit or checkout different versions as you see fit. This workflow would be more similar to a git workflow. But we encourage the use of Splitfiles because of the advantages they add; namely reproducibility due to provenance. It's sort of like how you can build a docker image by running arbitrary commands in a container and then `docker commit`. The problem with that workflow is that you lose all the benefits of Dockerfiles. The same logic applies to `sgr commit` and Splitfiles. Our bet is that data scientists will find Splitfiles to be the path of least resistance to accomplishing their goals. > where does query execution happen - always client or remote as well? At the moment, most of it happens on the client. But in Splitgraph Cloud, we do have the capability to execute queries on the remote. In a public setting, it's obviously more desirable to push down query execution to the client (or, if it's done remotely, to charge them for it). But in a corporate setting, you could imagine a shared remote cluster that executes queries on behalf of thin clients. So, it's possible to support both, but at the moment we're focused on the client. > in a global "Github for data" case is there some discovery mechanism for existing data Splitgraph Cloud includes discovery mechanisms including search and topics. We'll be adding a lot more features around this. We intend for the "data catalog" to be a core part of our offering. > do you provide a public storage to cover the case for Github for data? or is it now more like torrent - peers host and pay for data storage? At the moment, for simplicity and while we're in beta, Splitgraph Cloud is providing storage at our discretion. However, Splitgraph is designed so that data storage is decoupled from metadata storage. You can configure `sgr` to upload objects to any S3 compatible store. Currently it's configured to upload to object storage at Splitgraph Cloud, but there is no reason we could not introduce some kind of federation protocol where users can upload to independent silos of S3-compatible storage. But this raises a lot of questions with reliability and responsibility, so we have not fully explored it yet. In the near term, the easier solution will probably be charging clients for storage at Splitgraph Cloud. But, federation is something that is technically possible and at least academically interesting. Also, note that Splitgraph Cloud does not host all the data it includes in its index. For example, the 40,000+ datasets currently in the Splitgraph index are not hosted by Splitgraph [0], but we index them, and provide value added services like a REST API that does some remote execution of queries on your behalf. Currently these use the Socrata mount point, but you could imagine a situation in a corporate environment where the catalog might index lots of databases that are not Splitgraph images, but can be mounted with an FDW in the same way. > what is major direction for you - Github (public collaboration and sharing) or internal versioned warehouses (or some other internal case?). Most likely, both. We will probably follow the GitHub model of offering a public and on-premise version of the same product. In an ideal world, companies or universities might pay to license an on-premise version of Splitgraph Cloud that includes all the same features as the public version. We've done a lot of work on our backend to make deployments like this possible, so it's an appealing direction for us. [0] https://www.splitgraph.com/docs/splitgraph-cloud/external-repositories https://www.splitgraph.com/docs/splitgraph-cloud/external-re...