8 ms·
Modern Polars: A comparison of the Polars and Pandas dataframe libraries
- peepeepoopoo3 4y agoI hate pandas with a burning passion, but one thing it does have going for it is (some) interoperability with numpy, which opens up the rest of the scipy ecosystem. How easy is it to get numpy arrays into and out of polars?
- deleted 4y ago[deleted]
- culi 4y agowhat do you hate about pandas so much? I miss it dearly now that I don't use Python anymore
- Galanwe 4y agoPandas gets the job done, and is overall easy to use and intuitive. The problem is that it's a huge pile of hacks, exceptions, anti patterns, and regressions. The API is inconsistent, loose, full of obscure options added as quickfixes.
- Hasnep 4y agoI'm not GP, but I find the pandas API incredibly inconsistent and difficult to remember how to do simple transformations. For example, it sometimes overloads operators because it doesn't use built in language features like lambdas. There are reasons for the inconsistency, but using the alternatives like R's tidyverse or Julia's DataFramess.jl is like night and day for me. I found RedFrames [1] recently which wraps Pandas dataframes with a more consistent interface, it's probably what I'd use if I had to write data transformations that had to be compatible with Pandas. [1] https://github.com/maxhumber/redframes https://github.com/maxhumber/redframes
- pupperino 4y agoIt really can't be said enough how pandas is a mess. It has way too much surface area and no common thread pulling it all together. This gets obvious when you work with better dataframe libs like dplyr [1] or DataFramesMeta [2]. I've worked on production systems with all of these libs, this is not gratuitous bashing. [1] https://dplyr.tidyverse.org/ https://dplyr.tidyverse.org/ [2] https://juliadata.github.io/DataFramesMeta.jl/stable/ https://juliadata.github.io/DataFramesMeta.jl/stable/
- Hasnep 4y agoAn alternative I found recently is RedFrames [1] which wraps Pandas dataframes in a more consistent interface. That might be a better alternative if you need easy compatibility with Pandas. [1] https://github.com/maxhumber/redframes https://github.com/maxhumber/redframes
- fbdab103 4y agoThough that does look slick, the project is only ~5 months old. Which is a bit young for me to jump aboard.
- atoav 4y agoOh that looks interesting
- yuuuxt 4y agoSeems RedFrames is similar to pyjanitor, which is maturer if only comparing existence time: https://github.com/pyjanitor-devs/pyjanitor https://github.com/pyjanitor-devs/pyjanitor
- mbernstein 4y agoAs simple as a call foo.to_numpy() it looks like.
- ritchie46 4y agoVery easy. `pl.from_numpy` and `series.to_numpy` are your friend here. For 1D columns, we often can be zero copy as well. Besides that we support numpy ufuncs for `Series` and `Expressions`. As OP pointed out: https://kevinheavey.github.io/modern-polars/performance.html#numpy-can-make-polars-faster https://kevinheavey.github.io/modern-polars/performance.html... Numpy can be used to speed up some functions by utilizing numpy ufuncs. Numpy drops the GIL and therefore they can still be executed in parallel.
- topaz0 4y agoFunny seeing you here
- saeranv 4y agoDoes anyone else find the Polars syntax kind of clunky and ambiguous? For example, from the link, here's how Polars and Pandas handles manipulating data in a subset of a dataframe: f = pl.DataFrame({'a': [1,2,3,4,5], 'b':[10,20,30,40,50]}) # Polars f.with_column( pl.when(pl.col("a") <= 3) .then(pl.col("b") // 10) .otherwise(pl.col("b")) ) # Pandas f.loc[f['a'] <= 3, "b"] = f['b'] // 10 Its not clear in the Polars approach that the column "b" is being modified. An additional minor nitpick here is the use of when/then/otherwise for their conditional logic. Aren't these just if/else-if/else conditions? It's seems more in line with mathematical/python convention to use if/else... am I missing something? The Pandas equivalent, on the other hand, is much more concise, and more explicit. It also seems more mathematical to me. Polars mutates the dataframe, whereas in Pandas a function is applied to a dataframe indexed like a matrix. Pandas also benefits from it's reliance on symbolic notation, it makes everything visually clearer, whereas in Polars, the use of pl.col("b") and other similar methods contribute to multiple nested brackets and redundant naming calls contributing to less interpretability. I know there's a lot of thought thats been put into Polars, so I assume I'm missing some of the advantages of the Polars approach, and would appreciate anyone who can shed some light on it. I do understand, and partially agree, with the idea that indexing in Pandas leads to a lot of bugs. But in the example above, Pandas isn't really using indexing, it's using a boolean map to "index" the values from the same dataframe, so should be fairly robust. Is there a reason why Polars is trying to avoid this kind of filtering in the row/column indices?
- ritchie46 4y agoPolars author here. > Aren't these just if/else-if/else conditions? It's seems more in line with mathematical/python convention to use if/else... am I missing something? Yes, they are. But if you look at pandas `f['a'] <= 3` a boolean mask is created on eagerly, on the fly. Pandas has zero chance to do anything clever here. And yes, `when.then.otherwise` is exactly `if else`, but if `if else` is already a keyword in python so we cannot use them. `when, then, otherwise` are close synonyms. The benefit of using the `when().then().otherwise()` expression is that it is lazy. We don't do anything until we need to materialize the result. Then the optimizer has a chance to see the query a a whole and determine if the `mask` can be reused, is not needed, should be done somewhere else, etc. > Polars mutates the dataframe, Almost all polars methods are pure. There will be no dataframe mutated, but a new dataframe created. > Is there a reason why Polars is trying to avoid this kind of filtering in the row/column indices. Yes there is. Ambiguity. I want things to be explicit. So the method names should make clear that you are selecting rows: `df.filter` or selecting columns: `df.select` or slicing `df.slice` In pandas this can all be done with bracket notation. I often read code something like this `df[foo] = bar` and wondered what kind of datatype was stored into `foo`. Indexes has the same read complexity. I often read/saw queries that showed a different outcome after a `reset_index` call. I like things to be more explicit. This may cost some keystrokes, but future me/us can more easily understand what is going on.
- bfung 4y ago> Many of them are academic or quant types who seem to have some complex about being “bad at coding”. Glad I’m not the only one who’s noticed this. Coupled with this (which leads to: [I’m bad at coding so I won’t spend effort doing it even 1/2 way good]), and pandas having the most abstraction obfuscation of underlying data types, production can become a hot flaming mess that takes months to fix and scale up even linearly w/#of customers :sweat:
- stinos 4y agoGlad I’m not the only one who’s noticed this. Second. I understand that because of the the places I work I encounter this more than 'standard' (say web dev), but it's painful to see how much time and money this attitude seems to cost. Anecdotal rant incming, typical example encountered multiple times: person is really good at math but subpar at programming, but just enough to make it through a PHD (though I'm like 99% sure it's impossible there were no mistakes in that code). Anyway: pretty much every meeting the "I'm bad at programming" and "I don't really know anything about language/framework/thing X" is mentioned and used as if it's a valid excuse for messing up. But the worst part is: instead of just acting on it and learning and trying to improve, there's hardly any progress and without strict guidance anything touched by said persons turns into a trainwreck in no time. Again anecdotal, but I see this much less often with engineers.
- brahbrah 4y agoI have the same dynamic at my job. It’s a classic case of they can’t do what I can do and I can’t do what they can so let’s work together. It’s painful but necessary. I feel like it’s a perfectly valid excuse though. They have training in some other concepts that make them valuable, we can’t expect people to know what they know and learn what we know too. It’s why we work in teams. Although the people who are highly skilled in the analytical and engineering disciplines are worth their weight in gold.
- porker 4y ago> f.loc[f['a'] <= 3, "b"] = f['b'] Why isn't that saying to assign the value of column b to these locations? Reading the code (and not being a Pandas user) I expected it to be f.loc[f['a'] <= 3, "b"] = f['a'] Also the "// 10" comment is most confusing as looking at the result it matches 10, 20 & 30 in column b and replaces them with the matching values from column a
- antman 4y agoIf I understand correctly the currently promoted libraries for dataframes are: 1. Polars if data fits in ram 2. Vaex if data do not fit in ram 3. Spark with the dataframe api (koalas) if data do not fit in a computer Polars is great and delivers as promised
- pletnes 4y agoI actually thought polars’ lazy api would allow for out-of-ram computation? Also dask is more flexible than spark, since it lets you deal with numpy arrays and arbitrary objects better than spark can.
- ritchie46 4y agoIt does. Though the functionality is quite new, we will extend this. Calling `collect(streaming=True)` on a `LazyFrame` will allow you to process datasets that don't fit into memory. This currently works for groupbys, joins, many functions, filter etc. We will extend this to sorts and likely other operations as well.
- glogla 4y agoI'm curious if you could use this not for data science tasks but for data engineering tasks - say read a csv or pull a table from oracle and store it as delta lake table or something. I know its a boring use case, but the challenge with it is that it is a complete waste of money and carbon footprint to use Spark to process a 20 MB CSV or table with few thousand records, but tools like Pandas fall apart when you hit a 50 GB CSV or table with few billion records. Something more efficient (say, in Rust and not Python or Java) and yet scalable (due to not fitting everything into memory) would be a great help here.
- ritchie46 4y agoThis is exactly what we are aiming for. There are already a lot of queries that can be processed with 100s GBS of data on my 16GB laptop. And we will extend functionality for out of core processing. A single node can do a lot!
- bobbylarrybobby 4y agoA bit off topic but I would love to see the conciseness of the python polars API make it into rust. Mapping custom functions over a series is incredibly painful.
- Phlogi 4y agoIs there a way to download the whole book as a ebpub or kindle compatible document?
- __marvin_the__ 4y agoNot the author but it seems that the site was made using Quarto [1] which uses pandoc [2] behind the scenes for producing the final output. The pandoc website suggests EPUB is possible. [1] https://quarto.org/docs/get-started/authoring/text-editor.html https://quarto.org/docs/get-started/authoring/text-editor.ht... [2] https://pandoc.org/ https://pandoc.org/
- mutant_self 4y agoAuthor here: you are correct but the EPUB and PDF stuff actually didn’t work for this (it exited while rendering). iirc one of the problems was Quarto didn’t know how to format Polars dataframes
- sluijs 4y agoI could never get used to Pandas as a former user of R’s tidyverse. The naming and syntax never really sticked with me. I find Polars’ API much easier to reason about, and it definitely feels closer to dplyr than Pandas. I still miss the pipe operator though.
- lvass 4y agoThere's a great polars wrapper for Elixir called Explorer if you want pipes.
- mutant_self 4y agoThere’s a tidypolars package that appears to be well-maintained https://github.com/markfairbanks/tidypolars https://github.com/markfairbanks/tidypolars
- psimm 4y agoI feel the same. The closest to tidyverse in Python I've seen is siuba, a neat wrapper around pandas. Tidypolars is great too. Lately, I've used DuckDB to write SQL that manipulates pandas data frames.
- blakeburch 4y agoReally appreciate this side-by-side guide. Didn't realize Polars could still be used with Python, and the speed improvements seem to be drastic. May need to scope if it's worth updating our open-source connectors.
- erikcw 4y agoI just tried polars for this first time this week. I ported a data pipeline from pandas and I was blown away by the performance yield. Function went from a 60 min runtime with pandas to ~1:30 in polars! I’ve been using pandas for years and had no issues picking up the syntax. Can’t recommend giving it a try enough.
- brahbrah 4y agoBy any chance were you iterating over your pandas dataframe or using .apply? I’d be surprised by any properly formatted (i.e. vectorized) pandas operation that takes that long for data that fits in memory
- mutant_self 4y agoHere's an example of idiomatic Pandas taking 10 minutes while Polars takes 7 seconds: https://www.pola.rs/posts/the-expressions-api-in-polars-is-amazing/ https://www.pola.rs/posts/the-expressions-api-in-polars-is-a...
- brahbrah 4y agoI'm not saying that polars isn't faster. In fact in my other comment here I mention that polars is much better than pandas at what polars does (it's not a drop in replacement). I'm just saying that most of the times (not always, and in fact in those cases we've used polars to speed it up) that I've seen painfully slow pandas operations has been due to poorly formatted pandas code.
- brahbrah 4y agoI like polars a lot. It’s better than pandas at what it does. But it only accounts for a subset of functionality that pandas does. Indexes are not just some implementation details of dataframes. They are fundamental to the representation of data in a way where dimensional structure is relevant. Polars is great for cases where you want to work with data in “long” format, but that’s not always the most convenient way to work with data. Let’s say you want to get the difference in 15 day ahead temperature forecasts between forecasts on 2 different mark dates, for the forecast days they overlap (say the data consists of forecasted date, country, state equivalent, temp). In long format (necessarily in polars, optionally in pandas) you have to do: Merge df 1 and 2 on country, state and forecasted date, then create a new column of the diff between the 2 temp columns, then drop the 2 original temp columns. In a format where your indexes are forecasted dates on the rows and multiindex of country, state on the columns, you just have to do: df1 - df2 The way I see pandas is a toolkit that lets you easily convert between these 2 representations of data. You could argue that polars is better than pandas for working with data in long format, and that a library like xarray is better than pandas for working with data in the dimensionally relevant structure, but there is a lot of value in having both paradigms in one library with a unified api/ecosystem.
- n8henrie 4y agoI've been considering writing something like this for the last several days, glad the author has taken this off my plate :)