5 ms·
Ok, this was more complex than I thought. I hope I got it right and this gives you some pointers how I investigated. :) Use type() to see what you are getting:
by llao 8y ago
Ok, this was more complex than I thought. I hope I got it right and this gives you some pointers how I investigated. :)
Use type() to see what you are getting:
>>> type(df.describe().mean())
<class 'pandas.core.series.Series'>
>>> type(df.describe().mean)
<class 'method'>
Something without parenthesis cannot be a method call, at least not from a normal viewpoint. Instead you are trying to access the attribute "mean" of the result of describe() if you run df.groupby(df.var1).describe().mean. This attribute might very well be a method itself (a method object).
Last in that chain was describe(), so let's look at that: https://pandas.pydata.org/pandas-docs/stable/generated/pandas.DataFrame.describe.html https://pandas.pydata.org/pandas-docs/stable/generated/panda...
It returns a "summary: Series/DataFrame of summary statistics", so calling .mean() on that, will call https://pandas.pydata.org/pandas-docs/stable/generated/pandas.DataFrame.mean.html#pandas.DataFrame.mean https://pandas.pydata.org/pandas-docs/stable/generated/panda...
Accessing .mean will instead return that method object itself.
By default pandas will return a string describing the data itself when printing a method:
>>> df.describe()
numeric n2
count 3.0 3.0
mean 2.0 3.0
std 1.0 1.0
min 1.0 2.0
25% 1.5 2.5
50% 2.0 3.0
75% 2.5 3.5
max 3.0 4.0
>>> df.describe().mean
<bound method DataFrame.mean of numeric n2
count 3.0 3.0
mean 2.0 3.0
std 1.0 1.0
min 1.0 2.0
25% 1.5 2.5
50% 2.0 3.0
75% 2.5 3.5
max 3.0 4.0>
>>> df.describe().mean()
numeric 2.00
n2 2.75
dtype: float64
I assume this is what threw you off? You are not using the "results" of .mean in further code but were just looking at them, right?
- treyfitty 8y agoThat’s right, I was using notebooks for a project and luckily, didn’t have to actually extract the mean. Thanks for such an insightful answer
- llao 8y agoGlad to, as I need to learn Pandas in more depth than I currently know. Also, I noticed something very very dangerous you did: .describe() will return a DataFrame with all those statistics. One column per numeric column of the input. One row per statistical value (like mean, std, etc). If you call .mean() on THAT DataFrame, you will get the mean of the statistical values per column, so e.g. the mean of columnA's count, mean, std, etc values. Definitely not what you want! Instead, use either .mean() directly on your "data DataFrame" (not the one returned by .describe()). Or access the specific value from the result via .describe()['columnname']['mean'].