14 ms·
Polars: Fast DataFrame library for Rust and Python
- nas 5y agoIt looks interesting but phrases like "embarrassingly parallel execution" make my marketing hype detectors trigger. Maybe they could tone down their self promotion just a touch. Also "Even though Polars is completely written in Rust (no runtime overhead!) ...". I find that hard to believe.
- nojito 5y agoWhy? The benchmarks speak volumes. https://h2oai.github.io/db-benchmark/ https://h2oai.github.io/db-benchmark/
- sdfgsdf 5y agoThe benchmarks speak volumes of dishonesty. They sorted the results by speed of 1st run. For a language like Julia, which is JIT-compiled, that's not a fair comparison, considering that you compile once and run millions of times. Note also that Julia would be number 1 in almost all of those benchmarks if you were to rank by speed of second run (as expected...). It's funny because once you notice it those benchmarks are basically an ad for Julia. EDIT: Also..... lets think critically about some of the entries there. Most of them are languages, but then you have things like Arrow, which is a data format, Spark, which is an engine, ClickHouse and DuckDB are databases. The databases (and spark) will have to read from disk. They have no chance of competing with anything that's reading from ram, no matter how slow it is. They were built for different purposes. These are borderline meaningless comparisons.
- rscho 5y agoMaybe you should hop on the website of duckdb before commenting...
- paulgb 5y ago> considering that you compile once and run millions of times. If you’re writing data pipelines then yes, but a lot of Pandas users use it interactivity. As much as I’d rather use Julia, the last time I tried it I found myself waiting for computation far more often than with a Jupyter/Python workflow.
- deleted 5y ago[deleted]
- queuebert 5y agoGive it another try. They've improved the first run times quite a bit over the last few versions. Package precompilation has gotten way better as well.
- paulgb 5y agoGlad to hear it, I will!
- adgjlsfhk1 5y agoDataFrames1.3 is a lot faster specifically.
- kruxigt 5y agoWell, then it depends if interactively means redefined methods or just glueing together existing methods in new ways with new data. If it's the latter then no new compilations are needed. If it's the former then we are dealing with a type of work that can probably not be done in pure Python or similar and would require some kind of compilation anyway.
- nojito 5y ago>The benchmarks speak volumes of dishonesty. Not really. They are designed to showcase a common use case across multiple technologies. The beauty of this benchmark is that there is a hardware limit included so that it forces you to create novel solutions to perform well. >Note also that Julia would be number 1 in almost all of those benchmarks if you were to rank by speed of second run (as expected...). It's funny because once you notice it those benchmarks are basically an ad for Julia. Not sure where you're getting that but even on second run Julia doesn't really compete with DT/Polars
- prionassembly 5y agoJulia doesn't really compete with anything, despite having some cool tech behind it. It's like -- Julia is the Rory Gilmore of programming languages.
- adgjlsfhk1 5y agothe benchmarks are a bit out of date (missing DataFrames 1.2/1.3, Julia 1.7, CSV 0.9). I'm planning on running an updated version this weekend.
- 1egg0myegg0 5y agoIf you wouldn't mind, please update DuckDB as well!
- adgjlsfhk1 5y agoCan you make a PR to https://github.com/oscardssmith/db-benchmark https://github.com/oscardssmith/db-benchmark? I don't know DuckDB, so I don't know what the change would be.
- throwawaybutwhy 5y agoIt's obvious that you're promoting duck eggs at the expense of, say, chicken eggs or quail eggs or even ostrich eggs. Maybe you could tone that down a bit.
- apd_ 5y ago> Note also that Julia would be number 1 in almost all of those benchmarks if you were to rank by speed of second run (as expected...). Not true. If we'd rank them by second run Julia would be: - On simple query: 1st, 1st, 4th, 1st, 5th (down 1). - On advanced query: 3rd, 6th, 6th, 4th (up 1), - (out of memory). > The databases (and spark) will have to read from disk. They have no chance of competing with anything that's reading from ram, no matter how slow it is. Not true. Upon quick peek on the bench code, ClickHouse and Spark use in-memory table. I assume other engines too.
- sriku 5y agoAgree .. and I was looking for an option to sort by second run. One trick I've tried to some effect is to run jl code on a smaller data sizes so the compilation gets done and then repeat on the large one so it doesn't get interrupted by compilation. Not sure if this is a recommended approach. Benchmarking Julia is a pain for this reason - compilation always gets mixed up with runtime. But it hasn't prevented me from using it interactively. Pretty happy with it actually.
- ritchie46 5y agoNote that the compile times of julia are not included in the benchmarks. If you read the website, you'd seen that the grapsh show the first (excluding the compilation) and the second run (with hot cache). Also in the second run, julia is not the fastest. Julia would not be faster than Rust, its got a garbage collector. This is what you see in the join benchmarks that really push the allocator. Next to that, the databases run in in-memory mode, so there is not disk overhead. Spark is slower because JVM + row-wise data.
- rscho 5y ago> Julia would not be faster than Rust, its got a garbage collector. Having a garbage collector does not intrinsically make things slower. Especially so outside of the benchmarking microcosm.
- adgjlsfhk1 5y agothat said, Julia currently has a slow GC so it does hurt. GC performance is being worked on though. I have high hopes for a year or 2.
- sdfgsdf 5y ago> Note that the compile times of julia are not included in the benchmarks. If you read the website, you'd seen that the grapsh show the first (excluding the compilation) and the second run (with hot cache). Here's my view: The author of that page has commented here on HN; If my claim was so outrageously wrong as you claim, he would've corrected it.
- fault1 5y agoyeah, but your claim was "Note also that Julia would be number 1 in almost all of those benchmarks if you were to rank by speed of second run" notice this isn't even a language vs language benchmark. it's libraries and frameworks. plus I don't think even the author of the julia library in question would agree with your statement: https://discourse.julialang.org/t/the-state-of-dataframes-jl-h2o-benchmark/43081 https://discourse.julialang.org/t/the-state-of-dataframes-jl... as mentioned in that thread, GC and strings, or especially a combination of the two, can be very much a downer in terms of julia performance. That's actually pretty surprising since strings are often as important if not more important than numbers for a lot of data processing needs. I'd also say in terms of compilation time, some autocaching layer outside of precompilation would do wonders.
- space_rock 5y agoHe is basically describing benefits of the rest language so it's perfectly credible
- nas 5y agoHow so? Does Rust have zero runtime overhead? I would find that hard to believe.
- BenFrantzDale 5y agoIt’s a compiled optimized language. Along with C++, it’s one of the few languages to have essentially no runtime overhead.
- lern_too_spel 5y ago"Embarrassingly parallel" is a technical term, not a marketing term. https://en.wikipedia.org/wiki/Embarrassingly_parallel https://en.wikipedia.org/wiki/Embarrassingly_parallel
- nas 5y agoIt's a term for the nature of a problem, not a library or software package. It looks like they have designed the API so that "embarrassingly parallel" problems can naturally be computed using Polars. That would be fantastic, much better than Pandas. The way they write it sounds like marketing fluff to me and that's a shame because Polars looks like a useful thing.
- goodside 5y ago“Embarrassingly parallel execution” means that it parallelizes (only) problems that are embarrassingly parallel. The meaning is clear — if you want to be really pedantic about it, problems are “parallelizable” and only execution is “parallel”, but “embarrassingly parallelizable” is too many syllables.
- ritchie46 5y agoThe embarrassingly parallel is aimed at the expression API. This allows one to write multiple expressions, and all of them get executed parallel. (So embarrassingly, meaning they don't have to communicate and use locks).
- ahurmazda 5y agoI’ve been using it for the past quarter. In addition to the speed, I’m very pleased with the pyspark-esque api. This means migrating code from research to production is that much easier.
- jmakov 5y agoHow does compare to Vaex?
- VHRanger 5y agoThat's the real question
- rp1 5y agoThis question was asked last time the author posted this few months ago. I’m surprised they didn’t update the benchmarks. Kind of makes me think Vaex is faster.
- ritchie46 5y agoThe benchmarks are hosted by H2oAI, not by the polars team. Vaex is not in that benchmark. I don't believe Vaex would be faster though. They aim at larger than RAM data processing, not maximum in-memory performance like we do.
- abeppu 5y agoThere are so many dataframe libraries, many of which have APIs closely following pandas, but not drop-in replacements. I wish we could agree on a standard describing the core parts of what a dataframe must do, such that code depending only on those operations can easily move between dataframes.
- ddavis 5y agoThere is an effort for this: https://github.com/data-apis/dataframe-api https://github.com/data-apis/dataframe-api
- eternalban 5y agohttps://data-apis.org/dataframe-protocol/latest/design_requirements.html https://data-apis.org/dataframe-protocol/latest/design_requi...
- teruakohatu 5y agoWorse than that, pandas has a terrible API to start with. Going from the QueryVerse to Pandas feels like going back in time.
- bshipp 5y agoI've used pandas off and on for the better of a seven or eight years and, unlike other python libraries, I always feel like I'm starting from scratch every time I begin a new project. following a tutorial/official API on a second screen just to remember how to do fairly basic stuff. The reason is that once I'm done building whatever model I've needed it works so well I don't have to touch it again for a few years and I forget everything I learned (or the API changes again).
- sdfgsdf 5y agoIn Julia there's something better, called Tables.jl. It's not exactly an API for dataframes (what would be point the of that? You don't need many implementations of dataframes, you just need one great one). Instead it's an API for table-shaped data. Dataframes are containers for table-shaped data.
- xiaodai 5y agoIt's great to see innovation in this area.
- Maxion 5y agoI wouldn't really call it innovation, it's more just a project trying to bring to python something similar to the tidyverse from R.
- civilized 5y agoIn my world, anything that isn't "identical to R's dplyr API but faster" just isn't quite worth switching for. There's absolutely no contest: dplyr has the most productive API and that matters to me more than anything else. But I'm glad to see Polars moves away from the kludgey sprawl of the Pandas API towards the perfection of dplyr... while also being blazingly fast! Now just mix in a bit of DSL so people aren't obligated* to write lame boilerplate like "pandas.blahblah" or "polars.blahblah" just to reference a freaking column, and you're there! *If you like the boilerplate for "production robustness" or whatever, go wild, but analysts and scientists benefit from the option to write more concisely.
- extr 5y agodplyr API is not ideal in my experience. Overly verbose and confusing group/melt/cast operators. I much much prefer data.table. In your edit you mention concision, data.table is practically the platonic ideal of that!
- nuq 5y agoTrue that data.table is much simpler and faster one of the reasons I switched from dplyr to data.table
- civilized 5y agoMeh. Some people will never stop using Perl or APL because you can get anything done in five random characters (well, anything the language is optimized to express, everything else is a lot harder). I respect it but it's not for me. The tidyverse has the most advanced and intuitive versions of all the things you mention IMO. It has evolved a lot in the past couple years and your impressions of it could be out of date. There is also the dtplyr backend for data.table speed with dplyr syntax, but I don't even bother because dplyr is almost always fast enough for me.
- extr 5y agoI did go check out what's new in the tidyverse after your comment and was pleased to see new functions like pivot_wider and pivot_longer replacing the extremely confusing mess of spread and unite. So it's great to see the ecosystem evolving toward better usability. However I would hardly count it as a victory when late in the game you have to change the API for some core data manipulation functions because you made them too confusing the first time around. I think you are also maybe assuming everyone has the same use-case as you for data manipulation libraries. If you are coming from a non-programming context and picking up R for the first time, no doubt tidyverse is the way to do that. The verbosity is obviously a benefit if you're having to read someone else's code and are not interested in learning a DSL just to understand what columns are being filtered on or dropped or whatever. But if you are doing data analysis full time and are writing thousands of lines of throwaway EDA code a week, most of it only to be seen by yourself, the concision and speed that data.table offers is basically second to none, in any language. Rapid iteration for you personally is the point. Less typing is good, because you're trying to move as fast as possible to explore hypotheses. Execution speed on medium sized data is important, because a few extra seconds on every run matters a lot when you are running 500 micro-batches of analysis code a day. And as the h2o benchmarks show, data.table is still quite a bit faster than dplyr. Obviously not everyone needs the speed, but a lot of us do!
- Fiahil 5y ago… and it’s using arrow2, not the official, unsafe, arrow crate. Great, it means we can use it !
- vincent-toups 5y agoGod please anything to liberate me from pandas, which has one of the wildest API's I've ever had to routinely work with.
- rytill 5y agoHow would this compare to loading a sqlite database into memory and performing queries with it?
- 1egg0myegg0 5y agoPolars would be 10-100x faster, but so would DuckDB!
- rytill 5y agoWow, that’s amazing. I’ll definitely try it out. Do you know if there is any built-in functionality related to data compression or data loaders?
- riskneutral 5y agoI'm confused. Polars is built on top of the Rust of bindings for Apache Arrow. Arrow already has Python bindings. What does this project add by creating a new Python binding on top of the Rust binding?
- bogeholm 5y agoPolars is not using Rust bindings for Arrow, it uses a Rust implementation called arrow2: https://github.com/pola-rs/polars/blob/master/polars/polars-arrow/Cargo.toml#L12 https://github.com/pola-rs/polars/blob/master/polars/polars-... Arrow2: https://lib.rs/crates/arrow2 https://lib.rs/crates/arrow2
- thenipper 5y agoWe've been thinking about trying this out at work for some of our data pipelines/simplified models. The speed/ergonomics look great.
- sriku 5y agoHmmm .. in the linked benchmarks [1], DataFrames.jl (Julia library) appears to be fairly competitive. [1] https://h2oai.github.io/db-benchmark/ https://h2oai.github.io/db-benchmark/
- unixhero 5y agoWhat makes Pandas so bad and what makes Dplyr so great? I have used Pandas a lot for data analysis and for data integration duct tape scenarios. For me it has been a low bar for achieving a lot.
- StreamBright 5y agoI could never use Pandas without SO and the documentation and I use it for almost 10 years. I have no idea what is the intention of the developers most of the time.
- unixhero 5y agoAha, so you're productive right?
- bllguo 5y agoit's just so bloated and verbose. many ways to do the same things, annoying defaults (how is column not the default axis to drop?), indices are beyond frustrating (have never met anyone who doesn't just reset them after a groupby), inconvenient to do custom aggregations, very slow, not opinionated enough then there are the inherent python issues like dates and times, poor support for nonstandard evaluation, handling mixed data types and nulls
- wodenokoto 5y agoFor some people pandas seems to click. Good for you. I always struggle with google and the manual to get even simple things done. I can never figure out if I am gonna get a series or a data frame out of an operation. It seems to edit rows when I think it’ll edit columns and I constantly have to explicitly reset the index not to get into problems. I think dplyr is easy to read and write. It does get longer than other alternatives, but the readability is imho so good at it doesn’t feel verbose.
- otsaloma 5y agoIf you use Pandas daily, maybe get used to it and can ignore the issues, but for anyone using Pandas occasionally, it's every time a huge pain trying to figure out how to use it. The API is not intuitive and the documentation is very verbose and unclear. And stackoverflow top answers are often the "old way" of doing something when yet another way of doing the same thing has been added to the API.
- pvitz 5y agoDoes anybody here know dataframe systems that are able to handle file sizes bigger than the available RAM? Is polars able to handle this? I am only aware of disk.frame (diskframe.com), but don't know how well it performs.
- Fiahil 5y agoYou either stream them, or use bigger VMs.
- alexisread 5y agoI believe Vaex can do this, in addition to GPU processing and reading direct from s3. https://github.com/vaexio/vaex https://github.com/vaexio/vaex
- pvitz 5y agoTo you and all the other sibling comments: Thanks a lot! Exactly what I have been looking for! With regard to Vaex, I would really be interested in an independent benchmark comparing it to dask, spark, data.table etc. However, I have seen in the comments that others also can't find that.
- chrisaycock 5y agoThe H20 benchmarks cover Dataframe operations: https://h2oai.github.io/db-benchmark/ https://h2oai.github.io/db-benchmark/ It has pandas, dask, Spark, data.table, Polars, etc. Sadly, Vaex is currently missing from this suite.
- Matumio 5y agoFor Python there is Dask: https://docs.dask.org/en/stable/dataframe.html https://docs.dask.org/en/stable/dataframe.html
- cmollis 5y agospark dataframe api..
- 5y ago
- the_biot 5y agoI've never seen the term "dataframe" used as it is on this webste, and the commenters here seem to all use it. Judging by the examples it seems to just refer to a "row" from e.g. a CSV or SQL query. So is that all it is, or am I missing something?
- milliams 5y agoA "dataframe" is a "table"
- maxerickson 5y agoIt's a column oriented data structure.
- wodenokoto 5y agoA data frame is one of the basic, built-in data structures in R, which was released in 1993. And R was based on an even older S. So it’s not a new thing. If you don’t work in computational statistics / data science it might not be a well known term, though.
- gpderetta 5y agoFrom the python docs: > No Index > They are not needed. Not having them makes things easier. Convince me otherwise Agree completely. first class indices in pandas just complicate everything by having a specially blessed column that can't be manipulated consistently. Secondary indices should be "just" an optimization, while primary indices are a constraint on the whole table (not a single column). The library in general seem interesting. I'm not 100% sold on the syntax (as usual project is called select...), but it is not pandas which is already a huge plus.
- ritchie46 5y ago> (as usual project is called select...) Yeah.. this confusion is in the API as well (you can pass projection to IO readers). we used `select` because SQL. In the logical plan we make the correct distinction between selection and projection, but you don't see that very much in the API.
- Dowwie 5y agoPolars could bring the best of both worlds together if it can codegen python api calls to their Rust equivalent. A user conducts ad-hoc analysis and rapid development with Python. When the work is ready to ship, the user invokes a codegen to transform into Rust-equivalent api calls, resulting in a new rust module.
- callmerk 5y ago.
- ZeroGravitas 5y agoIs there a plugin to use this as a visidata backend? I quite like their UX.
- optimalonpaper 5y agoI'm reading all these comments and keep asking myself if I'm missing something, because I honestly sort of like pandas' API? Sure dplyr is nice -- it felt that way on rare occasions that I got to use it, at least -- but you get used to anything. So since I'm using python and know it quite well, I'm just more comfortable sticking with python's pandas framework rather than switching to R for data processing