9 ms·
Modern Pandas (Part 2): Method Chaining
- sterlinm 4y agoThis is a great series of articles but a bit funny to call it modern pandas these days since it’s six years old. Has it been updated?
- lysecret 4y agoFor large workflows, you'll probably want to move away from pandas to something more structured, like Airflow or Luigi. How are they a replacement for pandas ? I thought the would or at least could wrap around them for scheduled execution / chaining. You would still need a Data frame handling glibrart no?
- kzrdude 4y agoWhen the book is revised, I guess they should take all pandas features from 1.0 or some later 1.x release for granted.
- tomrod 4y agoTom Augspurger is one of those names you recognize when you subscribe to the pandas repo. He is very involved in the project! Great to see his thoughts and how he uses it.
- civilized 4y agoPandas may be OK for people who need to do a little data processing in their Python project, but I still recommend R and the tidyverse to those who need a serious tool of thought for analytics. Everything is so much more tidy and concise and intuitive and flexible. You can usually write code that directly expresses your high-level intent with a minimum of syntactic cruft. Full disclosure - the downside of modern R is the stack traces have been getting worse for a while.
- jmount 4y agoThere are a number of good packages in Python specializing in variations of powerful chained processing in Pandas. My own is this one: https://github.com/WinVector/data_algebra https://github.com/WinVector/data_algebra .
- civilized 4y agoI like the general direction of this. I do want to note that "there are many things for X in language Y" isn't necessarily a positive thing. It often means the community lacks the clarity of thinking or will to converge on one excellent product. Instead there are lots of okayish things each developed by a single person or handful of people.
- deleted 4y ago[deleted]
- d0mine 4y agoSignificant part of the work is massaging data and R sucks at that compared to Python. "if you want to do more than statistics, let’s say deployment and reproducibility, Python is a better choice." https://www.guru99.com/r-vs-python.html https://www.guru99.com/r-vs-python.html
- civilized 4y agoDepends on what kind of data. R is almost always better at tabular, Python is probably usually better at unstructured.
- beckingz 4y agoI was going to come here and post that I dislike method chaining because it is harder to read... And then I read TFA and realize that I actually dislike people writing poorly formatted method chaining pandas code. The examples in the post are really nicely formatted and easy to read!
- account-5 4y agoI've always found pandas really hard to use or reason about. I eventually get there but I don't like the code. Obviously subjective. I've never used another "data science" language though so I've no experience beyond it.
- mjburgess 4y agoEveryone does. Perhaps if python would add more support for FP (rather than its present hostility), we'd be able to phrase data transformations more naturally in the language.
- zasdffaa 4y agoMy python is rusty but IIRC it allows functional stuff (except for expression-only lambdas, boo). From https://docs.python.org/3/howto/functional.html https://docs.python.org/3/howto/functional.html it's got map/filter/currying and plenty more, what's misting in your view?
- mjburgess 4y agopattern matching expressions, syntax for partial application and composition, a typing system which can express structural types, a generalised list comprehension
- zasdffaa 4y ago> pattern matching expressions, nothing intrinsically 'functional' about this, though it's nice > syntax for partial application <https://docs.python.org/3/library/functools.html#functools.partial https://docs.python.org/3/library/functools.html#functools.p...> > and composition, Hmm, to my surprise I can't find this but TBH it's trivial to write, here's an eg <https://stackoverflow.com/questions/16739290/composing-functions-in-python https://stackoverflow.com/questions/16739290/composing-funct...> > a typing system which can express structural types, typing is orthogonal to FP (though very nice to have) > a generalised list comprehension IDK what this means, what's ungeneral about comprehensions currently?
- fumeux_fume 4y agoLove it. This is an excellent resource for brushing up on certain parts of the API I don't use as frequently. I've been using Pandas daily for about 6 years now and to see how it's evolved over time makes me really proud of the community. The obligatory comparisons to other tools in R are cliched and seem to completely miss the point.
- tda 4y agoWhat always bothers me is you can't method chain an in-place apply of some function on a (subset of) columns in an elegant way. Pipe and assign make it possible, but definitely not nice, just look at the example code I think I once proposed an apply_to method for this purpose but that got -1'nd by the creators in no time
- henrydark 4y agoQuick pdb trick for pipe-ers, you can stick this in the middle: .pipe(lambda df: (df, pdb.set_trace())[0])
- mayankkaizen 4y agoPandas is something that I wish I could avoid at any cost but I can't. There is simply no design philosophy. API is as ugly as it gets. I find it greatly unintuitive. It feels like a giant missmash of hacks on top of other hacks. Sometime I wish designers of Numpy or scikit-learn should have developed Pandas.
- patrick451 4y agoNumpy has plenty of API warts on its own. It's not obvious to me they would have done a better job at all.
- nuclearnice3 4y agoCould you elaborate on the warts?
- hervature 4y agoI always see these type of complaints and, when I actually sit down with people to resolve their aversion, it ultimately comes back to they use Pandas incorrectly or simply are not able to grok documentation. The usual dead giveaway is "Pandas documentation is horrible". Anyone who has used more than one documentation would know that documentation rarely includes every function in the API let alone the argument, examples, and links to other related functions. As this is the top comment, can you (and others) at least post the problems so that we can have an intellectual discussion? Maybe the Pandas devs might take a point or two.
- brundolf 4y agoIn my relatively brief experience maintaining a Python web service backed by Pandas: It's basically a DSL constructed out of the dismembered syntactic bones of Python, which breaks every piece of semantics in the host language that it possibly can. I'm sure this is convenient (and maybe even tractable to use) in an interactive notebook, where you can try out and verify behavior in real-time by looking at the output. But in any kind of non-immediate-feedback scenario where you're trying to engineer a production system, it's a hellscape of jumping back and forth to the documentation because even your most fundamental assumptions about the host language's semantics have been thrown out the window. On top of that (and largely because of it), it also resists static typing like crazy. It also has a bunch of APIs where it's really hard to know what will and won't mutate, and others that are just made generally hard to wrap your brain around for the sake of saving a few characters. Finally: some portion of the hate it gets is probably from engineers without a statistics background who aren't familiar with the huge dictionary of jargon and abbreviations its APIs use. This complaint is maybe less valid, since it is a data science library, but it certainly colors emotions and makes it even harder to deal with in a production context for lots of us. There were several times where, once I did some research and dug past the jargon and learned all the background for what it meant, the concept itself wasn't complicated. But Pandas did use the jargon, so I couldn't just understand, I had to go learn stats first and traverse all this needless indirection. It's an API that isn't designed for engineers but engineers often have to deal with it anyway.
- otsaloma 4y agoThe examples presented are probably carefully selected, my experience is that if you actually use Pandas with method chaining, (1) you get ugly-looking code due to a mixture of method calls and various different kinds of bracket indexing and (2) you eventually run into things that just can't be (nicely) chained and then you need to break the chain – even new stuff, such as DataFrame.append being deprecated in favor of pd.concat.
- _raoulcousins 4y agoOnce I discovered DuckDB, I use pandas methods a lot less. It’s so much easier for me to write SQL than to deal with Pandas’ syntax.
- jacobdi 4y agoMy team has been trying to modernize pandas from a different tact. Regardless of struggle with the syntax, it seems Pandas is very sticky, and we don't predict much migration to other data science languages. Instead of refining the syntax, we have combined it with a spreadsheet GUI (https://github.com/mito-ds/monorepo https://github.com/mito-ds/monorepo). Here, we worry less about writing perfect syntax ourselves, and let the GUI write the code for functions like pivot tables and merges that work well visually.
- lysecret 4y agoOk it seems I am the only person that loves pandas. And I have just recently started to like method chaining (even written some code to enable it with beautifulsoup) so this post came just at the right time.
- sterlinm 4y agoGiven how much of a role pandas seems to have played in the growth of Python over the last decade I suspect you aren't really the only person that loves pandas :) I think you'll find a similar selection bias if you ask HN commenters what they think about Excel.
- lysecret 4y agoGood point, also there are a lot of similarities between Excel and Pandas as well. I think this also is a fundamental distinction between normal SWEs and People who use Excel as well as Data Engineers. You always start with data, and you have no control over it. So this means: 1. you need state base programming env (excel, Jupyter) 2. you need to look at it to see whats there (plots) I guess HN is mostly comprised of SWEs who built the DBs and Websites that create the data that then gets consumed by the Data Engs and the Excel people :D
- sterlinm 4y agoAgreed about the similarities between Excel and Pandas. I started out as more of an data analyst and am now a SWE and I think one of the things that SWEs who dismiss Excel and Jupyter don't understand is how little you can assume about the data you might be working with. If you're an analyst who knows some VBA, that can be super useful but it would probably be a mistake to try to make your VBA-driven applications bullet proof. Nobody wants you to spend that much time on it, and the odds that something completely out of your control will change and break it anyway are quite high.
- MichaelRazum 4y agoI don't see how this ( df.pipe(went_up, 'hill') .pipe(fetch, 'water') .pipe(fell_down, 'jack') .pipe(broke, 'crown') .pipe(tumble_after, 'jill') ) is much better then something like that df = went_up(df, 'hill') df = fetch(df, 'water') df = fell_down(df, 'jack') df = broke(df, 'jack') df = tumble_after(df, 'jill') Would really like to hear an opinion about that.
- jeffwass 4y agoA former coworker of mine was a huge fan of functional programming, and also deeply allergic to mutation. So if you reused variables like that you’d get an angry earful. Though if you replaced each subsequent line with df1 and df2 and so on he wouldn’t mind as much. I can’t opine as to whether one approach or the other is intrinsically better. But echoes of his tirades still ring when I see the same variable name redefined.
- CuriousSkeptic 4y agoI would probably come across as matching that description. So perhaps I can speak for that position a bit. In this case I would not think re-assigning df as being in violation those principles. Might even be useful if it plays better with how the code interacts with debuggers and version control. Its not really re-used in the sense that would motivate making up new names for the intermediate steps. It’s clearly just used as a syntactic aid for the operation chaining. So in my mind both expressions are equivalent from that perspective. I would probably insist on limiting its scope to precisely that expression though, to maintain that obviousness.
- electroly 4y agoThere's no mutation there; it's just rebinding the name. Rebinding is very different from mutation. It wouldn't be my stylistic choice either but your FP friend wouldn't be complaining about mutation here.
- jeffwass 4y agoThanks for clarifying, though he definitely used the word “mutation” in this type of rebinding scenario.
- oneoff786 4y agoHmm I think this is bad. Not terrible, but bad. I’ll chain a few methods but never more than can easily fit on one line. Usually just something like groupby agg rename.
- qrian 4y agoThis article uses lambda function to change a column's dtype, but that is highly inefficient, and impractical once data size grows. Not sure it's because the article is relatively old (2016).
- tpoacher 4y agoYou don't need pandas to do chaining. It's a one-liner in pure python: https://github.com/tpapastylianou/chain-ops-python https://github.com/tpapastylianou/chain-ops-python Not to mention, it's a lot more debuggable this way (which is generally the biggest downside to most specialised chaining approaches).