3 ms·
(Airbyte Engineer) I think what you're saying here is often true until it isn't. For a personal project where you're pulling data from one API? Sure. Once you
by cgardens 6y ago
(Airbyte Engineer)
I think what you're saying here is often true until it isn't. For a personal project where you're pulling data from one API? Sure.
Once you have an engineering system with multiple engineers relying on the data to be pulled reliably, having a host of individual ELT crons gets brittle really fast. At the past couple companies I've worked out this same narrative has played out:
"Oh we need to pull data from X let's build a cron."
(3 months later.)
"Wait a second why is all of this data 1 month old? Oh the cron hasn't run in a month because of a schema change. Let's add monitoring."
(3 months later.)
"We need to change the cron to pull fields A,B,C hourly and field D,E,F weekly."
(3 months later.)
"The amount of data we're pulling is making this too expensive, we need to implement some sort of incremental replication."
... etc
It always starts out as a "small" script but they rarely stay that way. In my experience, they end up needing the same features that get rebuilt over and over again on an ad hoc basis. We want an engineer to be able to get these features out of the box.
We generally think that for most engineering teams (even pretty small ones) the ad hoc crons for pulling data become nightmarish pretty fast. This problem is compounded if you are already using some other ETL as a service provider but they don't support one of your data sources so you also have a separate set of crons. By taking an OSS approach we're trying to cover that long tail, so that all of your ELT can be managed using one tool.