4 ms·
I think there's a more common reason why companies end up with "awful-to-work-with messes": ETL is deceptively simple. Moving data from A to B and applying som
by romming 11y ago
I think there's a more common reason why companies end up with "awful-to-work-with messes": ETL is deceptively simple.
Moving data from A to B and applying some transformations on the way through seems like a straightforward engineering task. However, creating a system that is fault-tolerant, handles data source changes, surfaces errors in a meaningful way, requires little maintenance, etc. is hard. Getting to a level of abstraction where data scientists can build on top of it in a way that doesn't require development skills is harder.
I don't think most data engineers are mediocre or find their job boring. The expectation from management is that ETL doesn't require significant effort is unrealistic, and leads to a technology gap between developers and scientists that tends to be filled with ad-hoc scripting and poor processes.
Disclosure: I'm the founder of Etleap[1], where we're creating tools to make ETL better for data teams.
[1] http://etleap.com/ http://etleap.com/