4 ms·
I have tried polars for a couple of week and I am one of those weird guys who likes pandas syntax more than sql. Honestly, this is one of those "I will just wa
by anyfactor 4y ago
I have tried polars for a couple of week and I am one of those weird guys who likes pandas syntax more than sql. Honestly, this is one of those "I will just wait until it (Pandas) gets better" thing. Anyone who uses Pandas and SQL extensively knows that whatever question you might have with them someone already has an answer for you. One the other hand Polars is new and I feel like the Polars community is pushing the better syntax to wrangle data just doesn't feel right. I am not smart enough to put three different wrangling/query syntax in my brain.
I am hopeful about duckdb mostly because how friendly the people behind the project is. But honestly they really need to improve their csv reader operation. The data type recognition for auto read csv needs work. And duckdb people knows about CSV read is the reason polars has an edge on them and I know they are working on it.
But at the end of all that, I will just wait for Pandas to get faster and better.
- smohare 4y agoThis sounds like it’s spoken from someone who doesn’t understand a lick about sql. Pandas has one of the worst APIs I’ve ever seen. Truly a blight on the data processing landscape.
- brahbrah 4y agoIf you have a dataframe of power plant capacities and a dataframe of power plant capacity reductions how would you figure out the available capacities of all your power plants. In pandas it would be `capacities_df - reductions_df`. How would you do it in polars or sql, it’s not nearly as nice. Pandas has the benefit of allowing you to work with data in a relational or ndarray style. Rather than one or the other. The api does incur some bloat due to that, but that ability is very valuable.
- RandomWorker 4y agoHere here, I agree. In the most recent versions we have seen speed increases. The fact that polars exist shows that there is a tone of low hanging fruit. There is also dask to increase parallelism and performance which I’ve used on some massive datasets 200GB+
- __mharrison__ 4y agoPandas (the API) is also getting better at big data. I'm an advisor at a company, Ponder, that will take your Pandas code and execute it on "big data".
- tccole 4y agoAren’t there a bunch of plugins for that kind of thing?
- maegul 4y agoIsn't the low hanging fruit that polars picked: "how about lazy evaluation to allow the query to be optimised?" ... which is mostly anathema to the design of pandas?
- zeitlupe 4y agoI do not fully get the speed argument. If the dataset is small'ish, it does not significantly matter (given you only use the vectorized functions, of course). And if it's big data, I do not use pandas but Spark/a cloud data warehouse solution. For this reason I also do not get the use case for duckdb. Beyond this, a dataframe api/syntax, compared to SQL syntax, is to me way easier to follow and to debug.
- FridgeSeal 4y ago> I do not fully get the speed argument The way I see the speed improvements is this: fast processes run faster, moderate-to-long processes (that are still awkward and large in Python, but not enough to justify the shift to Spark) now run in noticeably faster time, and the threshold for “what I need to use Spark for, shifts a long way up the scale. That is, you can now do more, with the same amount of compute.
- orlp 4y agoThe use case for DuckDB is the that the vast majority of analytics aren't on "big data", but also aren't "small-ish". If your data fits on a single machine, DuckDB will allow you to query it at great speed. I disagree that the speed "doesn't significantly matter". There is work being done on a dataframe API frontend for DuckDB, if you prefer that interface.
- zeitlupe 4y agoSome more thought in the same direction: https://dataengineeringcentral.substack.com/p/whats-all-the-hype-with-duckdb https://dataengineeringcentral.substack.com/p/whats-all-the-...