8 ms·
I've had to dive into the pandas code over the last year for a project [0], and my attitude has shifted dramatically from... * old attitude: why does pandas
by closed 7y ago
I've had to dive into the pandas code over the last year for a project [0], and my attitude has shifted dramatically from...
* old attitude: why does pandas have to make things so hard
* new attitude: pandas has a crazy difficult job
I think this is most apparent in the functions that decide what "[d]type" a Block--the most basic thing that stores data in pandas--should be.
https://github.com/pandas-dev/pandas/blob/4edcc5541ff3f6470f5e3c083cb83136119e6f0c/pandas/core/internals/blocks.py#L2973 https://github.com/pandas-dev/pandas/blob/4edcc5541ff3f6470f...
And then, for the ubiquitous Object dtype, often figure out which of the many possible more specific types to cast it to.
If you think that is easy, ask yourself what this outputs:
import numpy as np
np.array([np.nan, 'a'])
Lo and behold--it produces an array where the np.nan has been converted to the string "nan".
And yet
import pandas as pd
pd.Series([np.nan, "a"])
Knows this, has your back, and does not stringify it.
It also has a pathological fixation on when it tries to convert dtypes, since avoiding all the bad conversion outcomes is a relatively time intensive process (compared to e.g. creating a numpy array).
I realize things could be much easier in pandas user facing interface, but really appreciate the sheer amount of effort that has gone into its dtype wrangling.
[0]: http://github.com/machow/siuba http://github.com/machow/siuba
- deleted 7y ago[deleted]
- _coveredInBees 7y agoI think another area pandas has done a lot of work on is with datetimes. Numpy's datetime objects are pretty deficient when you need to perform computations / data wrangling with them and utilizing python's native datetime objects would slow things down a decent bit. So they have done a lot of work to create their own datetime implementation that helps a lot when dealing with tabular data and performing date/time based arithmetic/manipulations.
- spectramax 7y agoDatetimes are still a mess. Now we have 3 datetime objects and converting between them is not obvious or trivial: https://i.stack.imgur.com/uiXQd.png https://i.stack.imgur.com/uiXQd.png
- _coveredInBees 7y agoYeah sure, but that isn't their fault really. I can see a role for python's datetime being separate from numpy/pandas, but I do think a consolidated datetime object would be better to have rather than the numpy and pandas versions that are similar but not the same.
- spectramax 7y agoDefinitely agree, pandas/numpy can have internal representations of efficient datetimes but the user should not have to deal with these conversions.
- logicchains 7y agoI use Julia and this is one area where Pandas still kicks the Julia ecosystem's proverbial ass: awesome support for working with nanosecond precision epoch timestamps.
- geoalchimista 7y agoThere are always choices to make. In this case, I would much prefer to let the data be treated as is, i.e., no silent casting of np.nan to the string "nan". When dealing with numerical data, a string "nan" is rarely useful. But when you need it, you can still create a data series with a string "nan" using pd.Series([str(np.nan), "a"])
- nl 7y agoSo you want the behavior that Pandas has (i.e., no silent casting of np.nan to the string "nan").
- geoalchimista 7y agoThe point I wanted to make was that string "nan" is usually not as useful as floating point NaN when dealing with numerical data. Therefore, as default behavior it is acceptable, if one has to make a choice.
- prepend 7y agoI like that pandas 1.0 took “Int64” out of trial mode. It wasn’t functionally very different from everything being a float, but I will appreciate not having to format floats as ints in all my reports.
- closed 7y agoAh, same! I'm also relieved to hopefully never again have to say: "whoops I converted nan into a very large integer".
- tel 7y agoI really, really dislike all the dtype wrangling and how those choices resonate throughout the API. I understand that a lot of work has been done to make that API "work", but in practice it feels like that effort would have been better avoided by changing expectations and interfaces. Now, to be clear, that's a hard problem. Heterogenous named bags of homogenous columns with a variety of data types, storage patterns, and ideas about missingness isn't an easy domain... but instead of just trying to make everything work through hammering 6+ semi-coherent interfaces (indices, databases, mutability, immutability/chaining, numpy, dataframes) together, I'd be willing to pay a lot more in verbosity and explicitness for something simple. pd.Series(str, [np.nan, "a"]) => ["nan", "a"] # or even an exception! pd.Series(nullable(str), [np.nan, "a"]) => [nan, "a"] Indexing is vastly over-designed. GroupBy is a very common API and is poorly documented and just weird in no small part due to attempts at dtype inference. Foundational useful concepts like categories feel bolted on. There's join, merge, pivot, pivot_table. I'd chalk this all up to just being "hard", but at the same time I can go pick up R's dplyr library and get a very nice existence proof of how a nice interface could work. Not to say dplyr has it all figured out, but it's a night-and-day improvement to Pandas. Pandas is great. It makes doing data science in Python so vastly much less of a chore than working with straight Numpy. It steals some great ideas and tries out a few interesting ones of its own... but it is far from a joy to work with.
- closed 7y agoAgreed RE GroupBy being challenging, especially compared to dplyr. As I've worked on a port of dplyr to python over the past year, though, I've realized the dtype issue (like you said), indexes, and GroupBy being difficult are likely connected. Basically, * dplyr can chop up a dataframe into 50,000 groups and apply arbitrary functions to it--no problem. * custom pandas grouped applies are very slow There are basically three reasons for slow pandas apply methods... 1. creating an index for each subgroup is slow (will not be a RangeIndex) 2. initializing a series for each group is slow (mostly due to type inference being re-run; could be avoided) 3. AFAIK more type inference is run when concatenating results This leads to a world where grouped calculations can't be run using arbitrary expressions (e.g. lambdas), but have to go through specific SeriesGroupBy methods. I wrote a bit on how I tried to work around that, to enable fast dplyr-like syntax over grouped data in python. Would definitely be interested in your take! There are other libraries, like ibis that do a good job with it, too! https://siuba.readthedocs.io/en/latest/developer/pandas-group-ops.html https://siuba.readthedocs.io/en/latest/developer/pandas-grou...
- ololobus 7y agoI dare to promote one StackOverflow question [1] about pandas I have tried to investigate and answer [2] half a year ago. And I was rather horrified by its internal complexity after digging into pandas source :) OP was wondering, why pandas facing a strange overhead after each 100th iteration in some very specific case. There was a proposal about Python's GC, but it was not clear at all. Finally, I have dived into pandas and found that it has a hard-coded constant == 100 (!) of a number of internal data storage blocks. After reaching this value it runs some consolidation routines [3], and they consume a lot of memory even leading to crash with memory error. What was much more wondering, is that after changing this constant to some large value (1000000, actually it disables consolidation at all) reduces memory consumption dramatically! This consolidation seems to reduce storage and memory consumption, so I still do not know why the opposite happens and why it works well in all other cases. [1] https://stackoverflow.com/questions/56690909/python-is-facing-an-overhead-every-98-executions https://stackoverflow.com/questions/56690909/python-is-facin... [2] https://stackoverflow.com/a/56705419/978424 https://stackoverflow.com/a/56705419/978424 [3] https://github.com/pandas-dev/pandas/blob/761bceb77d44aa63b71dda43ca46e8fd4b9d7422/pandas/core/internals/managers.py#L1171 https://github.com/pandas-dev/pandas/blob/761bceb77d44aa63b7...
- codedojo 7y agoHi, I'm working on this if anyone is interested (work in progress): https://github.com/dylan-profiler/visions https://github.com/dylan-profiler/visions The key idea is to allow abstraction of physical types (boolean, integers) by defining custom semantic types (URLs, paths, probabilities). The idea originated while working on pandas-profiling [0] and running into similar problems. We found this abstraction to be effective for many other downstream tasks, too, including compression and AutoML. More coming soon... [0]: https://github.com/pandas-profiling/pandas-profiling https://github.com/pandas-profiling/pandas-profiling