3 ms·
What are your thoughts on Pandas? As someone who uses numpy almost daily, I think that numpy is "overextended" beyond its core niche, sure. So - making it work
by Scene_Cast2 5y ago
What are your thoughts on Pandas?
As someone who uses numpy almost daily, I think that numpy is "overextended" beyond its core niche, sure. So - making it work with things outside that niche (e.g. streaming, non-rectangular data, non-uniform data, nonhomogeneous data, etc) is painful. However, 1) there's Pandas for that, and 2) I disagree with "misleading" and "surprising". What makes you think that?
- m_mueller 5y agoNot GP, but I’m using pandas daily to build up a BI platform within a financial institution. Compared to Matlab and even Fortran it has some issues IMO: * why distinguish between Series and DataFrame? just give me an interface for m x n matrices or even higher dimensions. * pure vs. in-place operations. not such a big fan of having multiple versions of the same function, e.g. a more pythonic df[“my_col”] = series vs. a more functional df.assign({“my_col”: series}) ; I’d rather have everything like the latter to be able to more easily have best practices in place. That brings me to another point: if we keep everything purely functional, then python’s syntax is making things a bit awkward. Where in something like JS you could just put every function call with its dot on a new line without the need to assign, in Python this requires putting line break characters or wrapping it in round brackets. This is one place where a language with explicit assignment terminators (semicolons) are a bit cleaner to work with. All that being said scipy is still a great choice to have both system programming and numerical business logic in one language.
- J253 5y agoIn my opinion and experience, I think you’re right about “there’s Pandas for that” and “that” can be almost anything. It can do almost anything but making it do almost anything requires constant reference to the docs. And I find maintainability difficult. It seems like there’s 50 kwargs for every method. Sometimes things happen in place by default, other times they don’t. Compound indexes still confuse me. But I’m not a data scientist so I don’t do much ad-hoc analysis that seems typical with pandas users.
- patrick451 5y agoNot the OP but I agree with them: Little things, like some functions want to be called with a tuple of dimensions, np.zeros((rows, cols)) others just want to be called like np.random.randn(n, m) The 1d array is a huge, fundamental design flaw in numpy. It makes zero sense that I can do matrix-vector multiplication against both an nx1 2d array as well as a 1d array. The latter is complete nonsense. When you slice a column from a matrix, and get a not an nx1 vector, but a 1d array, it makes me want to shell out $10,000 for matlab (yes, I know I can get a column vector with the slice A[:, [2]], but I shouldn't have to). This problem leaks out into the ecosystem. For example, when you try to use scipy to integrate an ODE, and pass it an initial condition vector that is nx1, the scipy integrator will silently coerce your vector to a 1d array, pass it to your RHS function, which then either blows up, or more likely, produces silently wrong result because of numpy's insane array broadcasting rules. This problem further leaks into the ridiculous function hstack. If you just used the function vstack, which made a 2x3 matrix from 2 1d 3 element arrays, you might imagine that hstack would produce a 3 x 2 matrix. But no. It creates a 1d 6 element array. For what you wanted, you actually need np.column_stack. I think the way Eigen handles this is the most intuitive. You do linear algebra with 2d objects, and cast to arrays for elementwise operations. There is also a huge inconsistency between what numpy exposes as an object oriented interface vs a "functional" interface. What I mean by this, is that I can call x.sum() on an array, but not x.diff(). For that, I need np.diff(x). There seems to be no pattern to what is exposed as a method vs a function. The array slicing api is also really inconsistent. For instance, given a 3 element array x, a = x[5] is an IndexError. However, this perfectly fine a = x[2:5] I just can't forgive that this is not also an IndexError.
- Yoric 5y agoI actually don't remember the details, I haven't used numpy in 4-5 years. I remember being bitten a few times by some operators that had a different behavior based on how you had arrived to what looked to be the same data. These were issues I don't remember encountering with e.g. Mathematica, MatLab or R, but then, I was manipulating different kinds of data. Next time I find myself manipulating numerical data, I'll definitely take a look at Pandas!
- civilized 5y agoPandas is better than nothing, but I would look to R's dplyr/tidyverse for a really well-designed tabular data manipulation ecosystem. Compared to tidyverse, the pandas API feels bloated, obscure, and inefficient. I often see people using very slow apply-based solutions in pandas because the faster solution is so non-obvious. The tidyverse ironically ends up feeling more Pythonic, with more of a "there is one obvious way to do it" vibe.
- da39a3ee 5y agoPandas is some probably very nice and clever Cython wrapped up in disastrous Python. As someone says below, doing anything requires constant reference to the docs, unless you did it yesterday. The semantics given originally to square bracket indexing have unacceptable edge cases with weird fallbacks, and instead of fixing it a bunch of other strange indexing syntaxes have been added on (but any Python programmer will use square brackets first). It's basically a distinct language (and a powerful one if you use it regularly).