6 ms·
It’s pandas, but fast. Pandas is the original open source data frame library. Pandas is robust and widely used, but sprawling and apparently slower than this ne
by DylanDmitri 3y ago
It’s pandas, but fast. Pandas is the original open source data frame library. Pandas is robust and widely used, but sprawling and apparently slower than this newcomer. The word “data frames” keys in people who have worked with them before.
- elmolino89 3y agoI may be a rare bird starting with R dataframes (still newbie+ level), then python polars (intermediate- ?). Frankly whenever I have to use pandas or df's in R I am not convinced that these are more intuitive/easier to master. I.e. I do not like the concept of row names. Polars can be an overkill for small/medium dataset, but since I have been bitten by corrupted/badly formatted CSVs/TSVs I love the fact that Polars will throw the towel & complain about types/column number mismatches etc. And the fact that it can scale up to millions of rows on a modest workstation compensates the fact that sometimes one can spend hours finding a proper way to manipulate a dataset.
- bee_rider 3y agoAh, like polar bears are a much more aggressive implementation of the idea behind panda bears? That’s a pretty funny name if so.
- Icathian 3y agoYeah. The name always makes me chuckle
- ayhanfuat 3y agoNot *original* but probably most commonly used.
- drbaba 3y agoYeah, I believe Pandas was inspired by similar functionality in R.
- froh 3y agoyup I first met data frames in R and pandas is the Python answer to R isn't it
- tomrod 3y agoIf I understand correctly, Pandas original scope was indexed in-memory data frames for use in high frequency trading, making use of the numpy library under the hood. At the time it was written you had JPMC's Athena, GS's platform, and several HFT internal systems (C++ my friends in that space have mentioned). Pandas just is so darn useful! I've been using it since maybe version 0.10, even got to contribute a tiny bit for the sas7bdat handling.
- froh 3y agoindeed it's both: it was created for financial analytics, and it provides R dataframe features to python. thanks for.making me detour into the history of it.
- gmfawcett 3y ago> Pandas is the original open source data frame library ...ehh, not quite. R and its predecessor S have Pandas beat by decades. Pandas wasn't even the first data frame library for Python. But it sure is popular now.
- p4ul 3y agoThat's interesting! I didn't realize there had been prior dataframe libraries in Python! Out of curiosity, what was/were the previous libraries?
- melagonster 3y agoit is built in data structure and function in R.
- p4ul 3y agoOh, yes, I was aware that R (and its predecessor S) have a native dataframe object in the language. It seemed that gmfawcett was indicating that there was a dataframe library in _Python_ that existed prior to Pandas. I was curious what that library was/is, as I'd not heard that before.
- melagonster 3y agook, guess I misunderstood both comments of you two. ´_>`
- gmfawcett 3y agoSorry :) Pandas is undisputed king. But there were multiple bindings from Python into R available in the early 2000's. Some like rpy and rpy2 are still around, others are long defunct. I concede that these weren't standalone dataframe libraries, but rather dataframe features built into a language binding.
- maliker 3y agoPandas has also moved to Apache Arrow as a backend [1], so it’s likely performance will be similar when comparing recent versions. But it’s great to have some friendly competition. [1] https://datapythonista.me/blog/pandas-20-and-the-arrow-revolution-part-i https://datapythonista.me/blog/pandas-20-and-the-arrow-revol...
- thejosh 3y agoMemory and CPU usage is still really high though.
- jasonjmcghee 3y agoNot according to DuckDB benchmarks. Not even close. https://duckdblabs.github.io/db-benchmark/ https://duckdblabs.github.io/db-benchmark/
- keithalewis 3y agoOuch! It is going to take a lot of work to get Polars this fast. If ever.
- hyperpl 3y agoPolars has an OLAP query engine so without any significant pandas overhaul, I highly doubt it will come close to polars in performance for many general case workloads.
- dash2 3y agoThis is a great chance to ELI5: what is an OLAP query engine and why does it make polars fast?
- disgruntledphd2 3y agoPolars can use lazy processing, where it collects all of the operations together and creates a graph of what needs to happen, while pandas executes everything upon calling of the code. Spark tended to do this and it makes complete sense for distributed setups, but apparently is still faster locally.
- dkga 3y agoActually pandas is not the original open source data frame library, perhaps only in Python. There is a very rich tradition in R on data.frames, which includes the unjustly neglected data.table.
- 7thaccount 3y agoYeah. I think Wes McKinney liked the data frames in R, but preferred the programming language of Python. I've heard somewhere that he also got a lot of inspiration from APL.
- Cacti 3y agoR is literally designed to do statistics and has first class support and language feature support for many specialized tasks in statistics and closely related fields. Python is literally designed to be easy to program with in general. Well, it turns out when you’re dealing with terabytes of data and TFLOPS, the programming becomes more important than the math. Not all R devs are happy about this and they are very loud about it. But it shouldn’t really surprise anyone. That is literally how those languages are designed. Most of the R devs I know like this are just butthurt they are paid less and refuse to switch because they’re obstinate, or they’re a little scared they’re being left behind. first group is all over the place, but the second group tends to skew older of course
- disgruntledphd2 3y ago> Most of the R devs I know like this are just butthurt they are paid less and refuse to switch because they’re obstinate, or they’re a little scared they’re being left behind. first group is all over the place, but the second group tends to skew older of course Look, I started with R and use mostly Python these days, but this is not really a fair take. R is (still) much, much, much better for analytics and graphing (the only decent plotting library in python is a ggplot clone). The big change (and why Python ended up winning) is that integrating R with other tools (like web stuff, for example) is harder than just using Python. pandas (for instance) is like an unholy clone of the worst features from both R and Python. Polars is pretty rocking, though (mostly because it clones from Spark/dplyr/linc). It's another example of Python being the second best language for everything winning out in the marketplace. That being said, if I was starting a data focused company and needed to pick a language, I'd almost certainly build all the DS focused stuff in R as it would be many many times quicker, as long as I didn't need to hire too many people.