4 ms·
Pandas has also moved to Apache Arrow as a backend [1], so it’s likely performance will be similar when comparing recent versions. But it’s great to have some f
by maliker 3y ago
Pandas has also moved to Apache Arrow as a backend [1], so it’s likely performance will be similar when comparing recent versions. But it’s great to have some friendly competition.
[1] https://datapythonista.me/blog/pandas-20-and-the-arrow-revolution-part-i https://datapythonista.me/blog/pandas-20-and-the-arrow-revol...
- thejosh 3y agoMemory and CPU usage is still really high though.
- jasonjmcghee 3y agoNot according to DuckDB benchmarks. Not even close. https://duckdblabs.github.io/db-benchmark/ https://duckdblabs.github.io/db-benchmark/
- keithalewis 3y agoOuch! It is going to take a lot of work to get Polars this fast. If ever.
- hyperpl 3y agoPolars has an OLAP query engine so without any significant pandas overhaul, I highly doubt it will come close to polars in performance for many general case workloads.
- dash2 3y agoThis is a great chance to ELI5: what is an OLAP query engine and why does it make polars fast?
- disgruntledphd2 3y agoPolars can use lazy processing, where it collects all of the operations together and creates a graph of what needs to happen, while pandas executes everything upon calling of the code. Spark tended to do this and it makes complete sense for distributed setups, but apparently is still faster locally.
- mjhay 3y agoLaziness in this context has huge advantages in reducing memory allocation. Many operations can be fused together, so there's less of a need to allocate huge intermediate data structures at every step.
- disgruntledphd2 3y agoyeah, totally, I can see that. I think that polars is the first library to do this locally, which is surprising if it has so many advantages.
- mjhay 3y agoIt's been around in R-land for a while with dplyr and its variety of backends (including Arrow, the same as Polars). Pandas is just an incredibly mediocre library in nearly all respects.
- disgruntledphd2 3y ago> It's been around in R-land for a while with dplyr and its variety of backends Only for SQL databases, so not really. Source: have been running dplyr since 2011.
- mjhay 3y agoThe Arrow backend does allow for lazy eval. https://arrow.apache.org/cookbook/r/manipulating-data---tables.html https://arrow.apache.org/cookbook/r/manipulating-data---tabl...
- vietvu 3y agoNot with eager API.