5 ms·
I agree in principle, but this is exactly why we've moved our data pipeline from Python to Haskell. The Python ecosystem has this concept of what is "pythonic"
by T-R 8y ago
I agree in principle, but this is exactly why we've moved our data pipeline from Python to Haskell.
The Python ecosystem has this concept of what is "pythonic", which, while it results in more code around the internet looking familiar, it largely resolves to "write it out yourself by hand". The result is that there are often no library functions for these things, even in libraries where you'd expect them to be. Need to coalesce an Optional, or need a switch statement? Write out the conditional. Need zipWith, flatMap, or mapMaybe? Write out the list comprehension. Need a combinator, or something curried? Write it out.
We spent an obnoxious amount of time fighting with bugs from this, or weird edge cases in pandas and numpy, that would show up halfway through training, or not at all until we measured results. Heisenbugs from generators were also a huge issue. We ultimately decided the situation wasn't tenable for the quick iteration cycles we needed, and moved all the data munging to Haskell. Almost everything we need for data munging is in base, usually in prelude. If it isn't, it's in lens. Changes and refactors are quick, and no more surprises halfway through training.
- kjeetgill 8y agoI like using zip/map/reduce in my data munge phases as much as anybody and I think python can become a ball of mud too fast sometimes. But your statement is just so out there to me: > it largely resolves to "write it out yourself by hand" me: Oh, I always felt the ecosystem was thorough between numpy, scipy, and bindings to most any database or C code like zookeeper, openCV, redis, etc. but I guess you needed something in a more specialized domain? > Need to coalesce an Optional, or need a switch statement? Write out the conditional. Need zipWith, flatMap, or mapMaybe? Write out the list comprehension. Need a combinator, or something curried? Write it out. ... I tried to be charitable, and maybe coming from a language with all of these things these are glaring omissions, but I had trouble no rolling my eyes. Whatever itertools doesn't have can be done in 2-3 lines. A little repetitive? sure. But so onerous it factors into language decisions? bizarre.
- T-R 8y agoIt's not about it being hard to write them out; everything that needs to get written out is surface area for bugs. One missed edge case could just mean that a node of your input layer gets zeroed for some of your rows where it shouldn't, and all you see is no correlation where you expect some. The little bugs that creep in writing this simple boilerplate kill days of work. One person writes unzip as zip(*arr), another writes it as two separate assignments with the list comprehension written out. It's a tiny bit of code; they both look like perfectly fine unzipping code, and they pass tests, but if you pass one of those code blocks a generator instead of a list - say, someone swaps a list comprehension for map - half your data disappears. No errors, just a node that shows no correlation. Since the results are all getting serialized anyway, the amount of work to just have that part of the pipeline in a language with some guard rails against that kind of thing is pretty minimal.
- stared 8y agoI agree that this is very annoying (and time consuming!) time for debugging, as it has implicit assumption about data (e.g. that some column has more than one value). I was tempted a few time to write pipelines in Python which make such sanity-checks. In any case - thanks for sharing your example. By any chance, can you show an example for such Haskell pipeline? (Especially if there is some non-trivial statistics, or ML training.)