4 ms·
Let me see if I can at least justify some of these things: - SQL became the de facto standard for data manipulation a few decades ago, and I’m still to see a w
by bobbruno 4y ago
Let me see if I can at least justify some of these things:
- SQL became the de facto standard for data manipulation a few decades ago, and I’m still to see a worthy contender. While the language itself shows its age, the concept of it being (almost) totally declarative of “what” data to get and “what” to do with it, instead of “how”, is unbeatable. A better syntax for achieving this is certainly possible, but you won’t beat decades of SQL code being generated easily. About “get the answer right once”, for analysis it’s often all that’s needed, and good SW maintence doesn’t apply that hard (it just won’t be needed again). For actual repeatable data engineering, I’d suggest either using generated SQL or writing in some abstraction (like PySpark). But don’t underestimate the productivity of SQL and the ability of modern parsers;
- Unit testing: while I generally agree, after 28 years, I have to say that unit testing is not as useful in data engineering. Essentially, it doesn’t matter how much time and effort you put on designing test cases, you’ll never beat the chaotic creativity of real users making real systems ingest all kinds of wrong data in patterns you’d never imagine. Now multiply that by having to cross-reference data from systems that were designed independently, and think of all the possible data errors that might come out of that. You won’t, you won’t cover even 20% of the ways things can go wrong. So, instead of “testing for all ways things go wrong”, a much better pattern is “write code that can capture unexpected errors and make reasonable decisions on whether to continue or abort - with some nice messaging and data debugging, please”;
- Poor support for version control: I 100% agree with you. DataOps is still not nearly as popular as it should be;
- Frameworks over libraries: let me tell you one thing - data is coupled. Uncoupled data has less value than coupled data, because it can’t answer more complex questions. Having different libraries that represent, manipulate and expose data in different ways will just make things much harder for the DE trying to generate the data. And the focus is handling the data, not the code. So, while you may have a bit of a point, I’d suggest that you’re looking at this from the wrong perspective;
Low-code: it’s one of those things that come and go. When I started, I’d use Oracle’s PL/SQL and C. Then came the first generation of visual ETL tools (Informatica, Datastage, etc) and they looked great with the graphic flows. Until you actually had to use them and all the clicking just to get to a point where you could actually add some logic showed. Then came the Hadoop era, and data engineering split into the DW people (who prefer visual, low-code to this day) and the Hadoop people, who went back to coding. This split stays to this day, with each side not recognising what the other does as essentially the same thing. Personally, I prefer code, but I don’t care what others prefer. The visual x coding discussion is, in my opinion, the wrong problem. The real issue is being able to create logic that’s clear enough that something as dumb as a computer will unambiguously understand. STEM learning environments for kids show that this is totally possible in a visual way - the thing is, most companies choose the visual tool in hopes that a person without the right mindset will generate good logic with it. That’s what fails. The rest is a matter of preferred UX - relevant, but not core or one-sided.