3 ms·
You made great points. Data Engineering cannot claim to support test-driven development to half the extent of the rest of software engineering. Frequently DE p
by vjust 3y ago
You made great points. Data Engineering cannot claim to support test-driven development to half the extent of the rest of software engineering.
Frequently DE projects have no concept of Dev, Test, Staging and then Prod. Quite often its dev and then straight to Prod. SQL of course is to partly or fully blame for this.
My last job for an large insurance company, they happily set a best practices of 90% test coverage (which in itself ended up being an artificial, ritualistic goal) which is impossible to achieve with DE tasks.
- atomicnumber3 3y agoIn my experience another big problem is that it's just expensive. For Spark jobs for instance, it's very common for it to run fine on the small test dataset on your laptop but then when you release it to run for 2 days on the prod dataset you end up having to do tuning there with pretty long turnaround times. And that then extends to why staging and even dev aren't very useful - the scale is part of the equation. Even just doing your dev loop, it's pretty easy to get something that's logically correct on your test sample but then when you run across the real dataset you find that 1 row out of every million is weird but you still gotta deal with it.
- necovek 3y agoYou seem to be touching on two disparate issues that I find very interesting to tackle. One is testing performance of code (SQL or otherwise) and ensuring it achieves a particular level of performance or ensuring no regressions. This is a problem not just with automatically unit-testing SQL, but any other code. In general, for regular code, we simply go with special, "manual" tests of the execution time for critical pieces (the biggest issue is fragility or flakiness: due to changing conditions a test runs under, speed is not always stable). We could do similar for testing SQL performance. If we know the database we are targetting, we could also use some of the introspection tools it offers to get even better tests (eg. we could run an "EXPLAIN (FORMAT JSON)" query on a very large Postgres database matching production to ensure right indexes are being hit and no seqscans are being done and Postgres' estimation of the time is on target). Basically, so far it's hard because "performance" means so many different things, but I don't think it's impossible. As for the other point, those 1-in-a-million edge cases, we've got those with regular code too! If you really want to test something against production-like DB, it's not hard (make an anonymised replica of prod DB), it's just expensive (tests will be slow, getting this set up will be slow, etc). I personally believe the right balance for cases like those is to catch them in production: automated tests should be quick and allow quick iteration if someone wants to do TDD on any part of the codebase. There are certainly product niches where this is not true (let's not have airbags in cars deploy accidentally every 1M rides, because there's a lot more than 1M rides daily :)), but for our regular applications, that's usually more than fine: quick tests will make it easy to fix the particular edge case once we hit it. FWIW, I love the idea of pgTAP for those who haven't seen it too.
- atomicnumber3 3y agoIt's probably worth noting that I was mostly talking about Data Engineering proper. I think "SQL is king" is (ime anyway?) comes from data analysts, who are usually not engineers are have more of a math of stats background. Data engineers do seem to write spark/pandas which is a different set of problems. I think a lot of the "big ball of sql mud" and brittleness just come from the fact that the people writing them aren't engineers and aren't following an engineering discipline. They're analysts trying to answer individual questions and find patterns to then pass off to ML engineering or data engineering teams to productionize. Or - they're lower-value pipelines that wouldn't be worth the time to ask an engineering team to try to prioritize them - so you either get a brittle thing that mostly works and need occasional love, or you do without it entirely.
- jochem9 3y agoWriting performing transformation code is one of the critical skills that a data engineer needs to master. It's a combination of experience with and knowledge about the underlying technology. How it actually processes data (e.g. learning how database pages work or going beyond dataframes and getting hands on experience with Spark RDDs).
- waffletower 3y agoData segregation requirements fundamentally break the utility of staged environments. There are some startups that try to fill this gap (Tonic.ai etc.) with data generation technologies -- yet it is extremely expensive to generate meaningful test data to populate development and staging environments that deeply mimic the interrelationships inherent to production data. These data relationships are very valuable to test. Because of this segregation, many organization resort to adhoc and manual testing techniques.