3 ms·
It's not about it being hard to write them out; everything that needs to get written out is surface area for bugs. One missed edge case could just mean that a n
by T-R 8y ago
It's not about it being hard to write them out; everything that needs to get written out is surface area for bugs. One missed edge case could just mean that a node of your input layer gets zeroed for some of your rows where it shouldn't, and all you see is no correlation where you expect some. The little bugs that creep in writing this simple boilerplate kill days of work.
One person writes unzip as zip(*arr), another writes it as two separate assignments with the list comprehension written out. It's a tiny bit of code; they both look like perfectly fine unzipping code, and they pass tests, but if you pass one of those code blocks a generator instead of a list - say, someone swaps a list comprehension for map - half your data disappears. No errors, just a node that shows no correlation.
Since the results are all getting serialized anyway, the amount of work to just have that part of the pipeline in a language with some guard rails against that kind of thing is pretty minimal.
- stared 8y agoI agree that this is very annoying (and time consuming!) time for debugging, as it has implicit assumption about data (e.g. that some column has more than one value). I was tempted a few time to write pipelines in Python which make such sanity-checks. In any case - thanks for sharing your example. By any chance, can you show an example for such Haskell pipeline? (Especially if there is some non-trivial statistics, or ML training.)