5 ms·
Some Python devs seem to pull in Pandas whenever any math is required. IMO Pandas documentation somehow manages to document every parameter of every method and
by fjp 7y ago
Some Python devs seem to pull in Pandas whenever any math is required.
IMO Pandas documentation somehow manages to document every parameter of every method and somehow it’s almost as helpful as no documentation at all. Combined with the fact that it’s a huge package, I avoid it unless I really really need it.
A version with human-understandable docs could convince me otherwise
- powowowow 7y agoI've found Pandas extremely easy to learn and to use; to the point where I find it confusing to see somebody say that it's not human-understandable. If you're reading this thread and wondering if it's easy to hard to use, I suggest taking a look at the docs (https://pandas.pydata.org/pandas-docs/stable/index.html https://pandas.pydata.org/pandas-docs/stable/index.html) and making your own decision. I find the combination of basic intros, user guides, and the API reference to be extremely usable and understandable; and I am reasonably sure I am human. But opinions may vary.
- anakaine 7y agoAnd I also. The other poster, in my opinion, has it backwards. Seeing how each option can be used in a myriad of examples is particularly helpful and saves me much time struggling to work out what the right method/ tool is, and how to use it.
- jfim 7y agoPandas has enough gotchas that it looks friendly until you hit one of them. Examples of gotchas: Want to join two dataframes together like you'd join two database tables? df.join(other=df2, on='some_column') does the wrong thing, silently, what you really wanted was df.merge(right=df2, on='some_column') Got a list of integers that you want to put into a dataframe? pd.DataFrame({'foo': [1,2,3]}) will do what you want. What if they're optional? pd.DataFrame({'foo': [1,2,3,None]}) will silently change your integers to floating point values. Enjoy debugging your joins (sorry, merges) with large integer values. Want to check if a dataframe is empty? Unlike lists or dicts, trying to turn a dataframe into a truth value will throw ValueError.
- qwhelan 7y ago>Want to join two dataframes together like you'd join two database tables? df.join(other=df2, on='some_column') does the wrong thing, silently, what you really wanted was df.merge(right=df2, on='some_column') Simply a matter of default type of join - join defaults to left while merge defaults to inner. They use the exact same internal join logic. >What if they're optional? pd.DataFrame({'foo': [1,2,3,None]}) will silently change your integers to floating point values. This was a long standing issue but is no longer true. >Want to check if a dataframe is empty? Unlike lists or dicts, trying to turn a dataframe into a truth value will throw ValueError. Those are 1D types where that's simple to reason about. It's not as straightforward in higher dimensions (what's the truth value of a (0, N) array?), which is why .empty exists
- jfim 7y ago> Simply a matter of default type of join - join defaults to left while merge defaults to inner. No, join does an index merge. For example, if you try to join with string keys, it'll throw an error (because strings and numeric indexes aren't compatible). left = pd.DataFrame({"abcd": ["a", "b", "c", "d"], "something": [1,2,3,4]}) right = pd.DataFrame({"abcd": ["d", "c", "a", "b"], "something_else": [4,3,1,2]}) left.join(other=right, on="abcd") ValueError: You are trying to merge on object and int64 columns. If you wish to proceed you should use pd.concat If you try to join with numeric keys: left = pd.DataFrame({"abcd": ["a", "b", "c", "d"], "something": [10,20,30,40]}) right = pd.DataFrame({"abcd": ["d", "c", "a", "b"], "something": [40,30,10,20]}) left.join(other=right, on="something", rsuffix="_r") abcd something abcd_r something_r 0 a 10 NaN NaN 1 b 20 NaN NaN 2 c 30 NaN NaN 3 d 40 NaN NaN Or even worse if your numeric values are within the range for indexes, which kind of looks right if you're not paying attention: left = pd.DataFrame({"abcd": ["a", "b", "c", "d"], "something": [1,2,3,4]}) right = pd.DataFrame({"abcd": ["d", "c", "a", "b"], "something": [4,3,1,2]}) left.join(other=right, on="something", rsuffix="_r") abcd something abcd_r something_r 0 a 1 c 3.0 1 b 2 a 1.0 2 c 3 b 2.0 3 d 4 NaN NaN Whereas merge does what one would expect: left.merge(right=right, on="something", suffixes=['', '_r']) abcd something abcd_r 0 a 10 a 1 b 20 b 2 c 30 c 3 d 40 d >> What if they're optional? pd.DataFrame({'foo': [1,2,3,None]}) will silently change your integers to floating point values. > This was a long standing issue but is no longer true. Occurs in pandas 0.25.1 (and the release notes for 0.25.2 and 0.25.3 don't mention such a change), so that would likely be still the case in the latest stable release. pd.DataFrame({"foo": [1,2,3,4,None,9223372036854775807]}) foo 0 1.000000e+00 1 2.000000e+00 2 3.000000e+00 3 4.000000e+00 4 NaN 5 9.223372e+18 It's also a lossy conversion if the integer values are large enough: df = pd.DataFrame({"foo": [1,2,3,4,None,9223372036854775807,9223372036854775806]}) foo 0 1.000000e+00 1 2.000000e+00 2 3.000000e+00 3 4.000000e+00 4 NaN 5 9.223372e+18 6 9.223372e+18 df["foo"].unique() array([1.00000000e+00, 2.00000000e+00, 3.00000000e+00, 4.00000000e+00, nan, 9.22337204e+18]) >> Want to check if a dataframe is empty? Unlike lists or dicts, trying to turn a dataframe into a truth value will throw ValueError. > Those are 1D types where that's simple to reason about. It's not as straightforward in higher dimensions (what's the truth value of a (0, N) array?), which is why .empty exists It's not very pythonic, though. A definition of "all dimensions greater than 0" would've been much less surprising.
- teej 7y agoYou can’t pick up a dictionary and determine how difficult it is to speak a language. One of the most difficult things about pandas is knowing how to shape your data frame before you plug it in to something. Most of my frustration comes from finding the magic incantation of symbols that will reshape my data frame the right way.
- Seufman 7y agoYeah, I agree: it's super straightforward. It's hard for me to understand how people could find it inscrutable / confusing. Pandas is a LIFESAVER for anyone who needs to manipulate datasets programmatically.
- mattrp 7y agoBig fan of pandas as well... but I have to admit there are some things at the beginning were like wtf... some of which are now deprecated thankfully. But more on topic to this post I really can’t see why someone would want to solve speed on small datasets by incorporating numpy into a new form of pandas... both projects are so established, why would you attach your project to some other dev who know has to keep pace with numpy and pandas improvements when you could just import pandas, import numpy and be done with it?
- edgyquant 7y agoI learned Python by using pandas so for awhile I was one of the devs you speak of. I can't imagine I'm alone.
- missosoup 7y agoPandas has some of the best documentation that I've seen. Maybe the issue is that the documentation makes a lot less sense without a stats and data background, but expecting pandas documentation to be a stats course is silly. Other thing is both numpy and pandas allow for some pretty complex operations by chaining simple commands. The space of what's possible is effectively unbounded, documentation won't help there. Cookbooks cover some of the more common patterns.
- acomjean 7y agoI feel like the documentation is great if you already know how it works. Having learned Pandas recently, I'll confess there are parts that are a bit tricky to get past the basics. I know R dataframes which was helpful. Pandas is well worth it to learn however if your using python.
- cerved 7y agoI imagine OP is referring to things like conditional selection on multiple columns, not by AND && operators but by bitwise & operators. I've mostly use comprehensions to filter and manipulate data and I find the Panda API to be a bit clunky and esoteric. There are nice built in functions and it displays well in a jupyter notebook but I'm not a fan of the interface.
- missosoup 7y agoI mean yes, it's a bit esoteric, but so is every other data manipulation package/language. Look at numpy, dplyr, q, spark, or even sql for similar use cases. There's no easymode way to express arbitrary data operations, especially not when you want them to be performant. If you don't need the advanced data manipulation capabilities of tools like pandas/dplyr (and if you're using comprehensions only, you don't), there are much easier to use options like apache nifi. Horses for courses.
- Grimm1 7y agoPandas' type conversions irk me. I get why columns with NaN convert to float from integer but very rarely do I have data that is complete for every column and converting columns that were intentionally integer has caused headaches when that data then goes to other systems such as a sql db. It is currently sitting at the center of an ETL system at my work (not my decision) and causes headaches.
- abakker 7y agoI mean, whenever I have this problem, I fall back on my (terrible) SPSS habits and just recode NaN to -999. With integer datasets that mostly works fine. You could come up with some alternative solution, too. Personally, I'm an amateur at best, but Pandas has made it possible for me to do some really handy things. Personal favorite feature: multindexing on rows AND Columns. Multiindex is such a common pattern, and so poorly handled in things like excel. Pandas really saves a lot of time with re-indexing, pivoting, stacking, or transposing data with multi indexes.
- qwhelan 7y agoYou should upgrade your version of pandas if possible - that's been fixed for a few versions now.
- Grimm1 7y agoOh interesting I'll have to check, we did an upgrade pass a while ago maybe we just didn't upgrade pandas for whatever reason.
- qwhelan 7y agoAs mentioned elsewhere in this thread, it's opt-in to avoid breaking existing behavior. But given that ingestion points are easy to identify, it's pretty straightforward to turn on (especially if you have a schema for your inputs): https://pandas.pydata.org/pandas-docs/stable/user_guide/integer_na.html https://pandas.pydata.org/pandas-docs/stable/user_guide/inte...
- mhh__ 7y agoI end up using it because it's there for the most part. I use python for hacking stuff, and not because it's fast but because the libraries exist. I am almost repulsed by the language and now you've said it I agree wrt the documentation too.