4 ms·
1. I think in practice, people are already using a mix of sources for ingestion today. It’s rare that a data team would rely on a single tool to ingest all thei
by ctc24 4y ago
1. I think in practice, people are already using a mix of sources for ingestion today. It’s rare that a data team would rely on a single tool to ingest all their data – instead, they might get some data via one or more ETL tools, some data via a script they wrote themselves, and some other data from their own db. So in that regard, I don’t think a world where the vendor provides the data pipeline makes this a lot more complex.
One way we’re hoping to pre-empt some of this is by helping vendors to surface more observability primitives in the schema that they write data to. To give an example: Prequel writes a _transfer_status table in each destination with some metadata about the last time data was updated. The goal there is to decouple the means of moving data with the observability piece.
We can also help vendors expose hooks that people’s data pipelines & observability tools plug into (think webhooks and the like).
2. We don’t really – anecdotally, companies that offer data warehouse syncs tend to be pretty focused on providing a great user-experience. At the end of the day, that decision is pretty much entirely with the business. We see a pretty wide range today: some teams choose to make exports widely available, some choose to reserve it for their pro or enterprise tier, and some choose to sell it as a standalone SKU. It’s pretty similar to what already happens with APIs designed for data exports.
Would love to continue the convo and hear more about your take on obs! Sending you a note now.