3 ms·
I can't say this is going to make a big difference in how I use pandas but I've ran into the bizarre "can't have nans in an int Series" annoyance in almost ever
by mactrey 8y ago
I can't say this is going to make a big difference in how I use pandas but I've ran into the bizarre "can't have nans in an int Series" annoyance in almost every pandas project I've worked on, so good on them for fixing that.
- em500 8y agoIt might be annoying, but it's certainly not bizarre. In the floating point standard (IEEE 754) there are reserved bit values (with hardware support) to represent NaNs. For integers no such thing exists, so you're left with a bunch of different implementation choices all with different tradeoffs. A long time ago NumPy devs choose not to support NaN in integer arrays at all (for maximum performance), and Pandas (starting as a wrapper around numpy arrays) inherited that. For a more technical discussion of the performance implications of one possible implementation, see http://wesmckinney.com/blog/bitmaps-vs-sentinel-values/ http://wesmckinney.com/blog/bitmaps-vs-sentinel-values/
- mFixman 8y agoThere's a nice explanation for this in Pandas' FAQ: https://pandas.pydata.org/pandas-docs/stable/user_guide/gotchas.html#support-for-integer-na https://pandas.pydata.org/pandas-docs/stable/user_guide/gotc... Not allowing nullable type in raw integer `Series` is the least bad solution this problem. If you really want nulls in numerical data, either use floating point or `IntegerArray`.