2 ms·
I've been making these comparisons between the new dataframe libraries in python so a couple of comments: "Reading the full CSV without datetime parsing is in
by braaannigan 4y ago
I've been making these comparisons between the new dataframe libraries in python so a couple of comments:
"Reading the full CSV without datetime parsing is in line in terms of speed though."
This sentence was a bit ambiguous, but is important: if you read this file in pandas with engine='pyarrow' but don't convert the date/time column to a pandas datetime dtype you get the same ~100 ms read time as calling PyArrow directly. So basically the entire time is spent converting the strings to dates. If this datetime thing isn't an issue for you then you can just use the engine='pyarrow' argument to read CSVs with Pandas.
In my own tests with various datesets Polars has always been much faster than duckdb/pyarrow. For this relatively small dataset it's about 2x faster, which is about the smallest margin I've found. Polars is also much easier to write, as its query optimization is so effective - you don't need to know all the tricks that Pandas requires (and is still 3x-10x faster than Pandas even when you apply all the tricks).
I've started making videos to address the need for more guidance in Polars - see this new on reading CSVs: https://www.youtube.com/watch?v=nGritAo-71o https://www.youtube.com/watch?v=nGritAo-71o