7 ms·
Nice article. Pandas gets the job done but it's such a step backwards in terms of useability, API consistency and code feel. You can do anything that you can po
by beforeolives 5y ago
Nice article. Pandas gets the job done but it's such a step backwards in terms of useability, API consistency and code feel. You can do anything that you can possibly need with it but you regularly have to look up things that you've looked up before because the different parts of the library are patched up together and don't work consistenly in an intuitive way. And then you end up with long lines of().chained().['expressions'].like_this(0).
- nerdponx 5y agoI much prefer writing actual functions in a real programming language to the verbose non-reusable non-composable relic that is SQL syntax. Give me a better query language and I will gladly drop Pandas and Data.Table.
- beforeolives 5y agoYes, SQL does have the problem of bad composeability and some things being less explicit than they are when working with dataframes. Dplyr is probably the most sane API out of all of them.
- alexilliamson 5y ago+1 for dplyr. I've used both pandas and dplyr daily for a couple years at different times in my career, and there is no comparison in mind when it comes to usability/verbosity/number of times I need to look at the documentation.
- RA_Fisher 5y agoI agree, the degree of expressiveness that dplyr achieves is impressive. The library has had a great impact on my work. The syntax is easy to remember so you don't have to constantly look up minor variations in usage. Also, the pipelining aspect makes it dead-simple to debug. It's especially powerful when combined with the purrr library: https://purrr.tidyverse.org/ https://purrr.tidyverse.org/
- semitones 5y agoSQL is a real programming language.
- thinkharderdev 5y agoIt's not turing complete which is I think why it doesn't feel like a "real" programming language.
- sagarm 5y agoRecursive CTEs exist; while I haven't investigated thoroughly, I would be surprised if they didn't make SQL turing complete. More usefully, most databases allow you to write user-defined functions in other languages, including imperative SQL dialects.
- steve-chavez 5y agoSQL is turing complete. A demonstration with recursive CTEs is done here[1]. [1]: https://wiki.postgresql.org/wiki/Cyclic_Tag_System https://wiki.postgresql.org/wiki/Cyclic_Tag_System
- thinkharderdev 5y agoOh wow. I stand corrected. Thanks!
- nvilcins 5y agoYou might be thinking of specific functionality that you find is being implemented in an overly long/verbose fashion.. But generally speaking, how are > long lines of().chained().['expressions'].like_this(0) a _bad thing_? IMHO these pandas chains are easy to read and communicate quite clearly what's being done. If anything, I've found that in my day-to-day while reading pandas I parse the meaning of those chains at least as efficiently as from comments of any level of specificity, or from what other languages (that I have had experience with) would've looked like.
- musingsole 5y agoPeople don't like them because the information density of pandas chains is soooo much higher than the rest of the surrounding code. So, they're reading along at a happy place, consuming a few concepts per statement ...and then BOOM, pandas chain! One statement containing 29+ concepts and their implications. Followed by more low density code. The rollercoaster leads to complaints because it feels harder. Not because of any actual change in difficulty. /that's my current working theory, anyway
- disgruntledphd2 5y agoHmmm, interesting. I don't mind the information density, coming from R which is even more terse, but the API itself is just not that well thought out (which is fair enough, he was learning as he went).
- isoprophlex 5y agoThis really grinds my gears too. There's something about the pandas API that makes it impossible for me to do basic ops without tedious manual browsing to get inplace or index arguments right... assignments and conditionals are needlessly verbose too. Pyspark on the other hand just sticks in my brain, somehow. Chained pyspark method calls looks much neater.
- ziml77 5y agoSetting inplace=True isn't too bad, but I definitely have had many issues with working with indexes in Pandas. I don't understand why they didn't design it so that the index can be referenced like any other columns. It overcomplicates things like having to know the subtle difference between join() and merge().
- nojito 5y agoSetting Inplace=True is not recommended and should be used with caution https://github.com/pandas-dev/pandas/issues/16529 https://github.com/pandas-dev/pandas/issues/16529
- ziml77 5y agoI've never seen anything about inplace being a poor idea to use. Is that documented anywhere or is that ticket the only info about it?
- Peritract 5y agoIt's not particularly front-and-centre, but it's all over various discussion boards if you go looking for it directly. I'd like the debate to be more visible, personally, particularly whenever the argument gets deprecated for a particular method. Essentially, `inplace=True` rarely actually saves memory, and causes problems if you like chaining things together. The people who maintain the library/populate the discussion boards are generally pro-chaining, so `inplace` is slowly and quietly on its way out.
- carabiner 5y agoCan you honestly say you'd prefer to be debug thousands of lines of SQL versus the usual <100 lines of Python/pandas that does the same thing? It's no contest. A histogram in pandas: df.col.hist(). Unique values: df.col.value_counts(). Off the top of your head, what is the cleanest way of doing this in SQL and how does it compare? How anyone can say that SQL is objectively better (in readability, density, any metric) other than the fact that they learned it first and now are frustrated that they have to learn another new tool baffles me. I learned Pandas first. I have no issue with indexing, different ways of referencing cells, modifying individual rows and columns, numerous ways of slicing and dicing. It gets a little sprawling but there's a method to the madness. I can come back to it months later and easily debug. With SQL, it's just madness and 10x more verbose.
- beforeolives 5y ago> Can you honestly say you'd prefer to be debug thousands of lines of SQL versus the usual <100 lines of Python/pandas that does the same thing? That's fair, I was using the opportunity to complain about pandas and didn't point out all the problems SQL has for some tasks. What I really want is a dataframe library for Python that's designed in a more sensible way than pandas.
- marktl 5y agoDistinct Count and Group By in SQL will provide distinct Count and Histogram (numerical value) output. Like No SQL vs Relationship, seems there are use cases that lemme them selves to new vs old technologies. It's hard for me to put much stock in opinions from those that have vastly more experience in one tool and not the other being compared.
- bayan1234 5y agoThis is exactly my experience. I've found becoming proficient with Pandas to take much more time than other data manipulation libraries like dplyr or Pyspark dataframes. The Pandas way of doing things is just not memorable or intuitive.