3 ms·
Hey George, thanks for the comment and for the good points! We want to initially focus on the use case with lots of diverse backend datasets and ad-hoc APIs (m
by mildbyte 6y ago
Hey George, thanks for the comment and for the good points!
We want to initially focus on the use case with lots of diverse backend datasets and ad-hoc APIs (maybe with a no-code like solution on top of a generic FDW) where performance won't be the bottleneck. If necessary, the backend data sources can perform aggregation and fast query execution. For example, you can also put Splitgraph in front of Presto (through JDBC). The value we want to provide in these cases is:
* granular access control (e.g. masking for PII columns, auditing etc)
* firewalling/query rewriting/rate limiting (for publicly accessible endpoints that proxy to internal databases that vendors want to publish more easily than through cronjobs with data dumps)
* cataloguing (so you get to discover datasets/data silos, get their metadata and query it over multiple interfaces in the same product)
We also like keeping the PG wire format in any case, as there are so many BI tools and clients that use it that it makes sense to not break that abstraction. We started with PG FDWs just because of the simplicity and the availability of FDWs, but we might swap the actual Postgres FDW layer for some faster execution in the future, if it's needed.
Behind the scenes, for Splitgraph images, we use cstore_fdw as an intermediate storage format (it's a columnar store similar to ORC with support for all PG types like PostGIS geodata). There's a potential in using this as a format for a local query/table cache on Splitgraph nodes that we intend to deploy around the world for low-latency read-only query execution.