3 ms·
Its not trivial to create and manage Data Pipelines if you care about scale, serving a wide range of inputs and outputs, or making this data easy to surface and
by prions 6y ago
Its not trivial to create and manage Data Pipelines if you care about scale, serving a wide range of inputs and outputs, or making this data easy to surface and spread throughout your org (i.e. making it actually useful to regular people).
"Static ETL" like running the same database load every day at 1:00am isn't a super challenging problem.
Doing it across many tables with complex transformations and multiple steps easily can be. You really have to consider reliability, processing speed, failure methods and other problems that dont really arise until you hit a certain scale.
There's also the issue of what people want out of a Pipeline that's changing. If you want people to be ""data driven"", then that means they need easy access to potentially all of your company's data on an ad hoc basis. So now your boring ETL 1 am pipeline isnt really serving any of these new usecases.
How do you create flexible pipelines that can be created from any dataset on an ad hoc basis? This is where tools like Airflow or Prefect come in. Creating a platform that can create these types of Pipelines is a real problem.
And before you even ask yourself _how_ to process this data, you need to also ask _where_? If you want to do what I outlined above - making your data more accessible and easy to use - then you probably need to rework how you're storing your data. But Data Lakes (and others) are a whole topic in and of itself.