4 ms·
Author of polars here. If we would have included csv parsing in the benchmark, polars would have come out much better than it already has.
by ritchie46 4y ago
Author of polars here. If we would have included csv parsing in the benchmark, polars would have come out much better than it already has.
- henrydark 4y agoBetter than the apache arrow c++ implementation that's in pyarrow?
- ritchie46 4y agoIt is on par. Just did a local benchmark on the new york taxi dataset. ``` # polars We should not rechunk as that is not what pyarrow does. Rechunking is a different operation and is also optimized away in many lazy operations. %%time pl.read_csv("csv-benchmark/yellow_tripdata_2010-01.csv", rechunk=False, ignore_errors=True) CPU times: user 17.4 s, sys: 2.42 s, total: 19.8 s Wall time: 1.89 s %%time pa.csv.read_csv("csv-benchmark/yellow_tripdata_2010-01.csv") CPU times: user 17.1 s, sys: 3.07 s, total: 20.1 s Wall time: 1.99 s # pyarrow embedded new lines (e.g. valid csv files) Polars by default allows embedded new line characters. Pyarrow does not, if we tell it to do so, polars is faster on my laptop. %%time pa.csv.read_csv("csv-benchmark/yellow_tripdata_2010-01.csv", parse_options=pa.csv.ParseOptions(newlines_in_values=True)) CPU times: user 17.8 s, sys: 3.09 s, total: 20.9 s Wall time: 2.72 s # pandas (pyarrow engine) And pandas itself has to pay for the copy from pyarrow to pandas: %%time pd.read_csv("csv-benchmark/yellow_tripdata_2010-01.csv", engine="pyarrow") CPU times: user 18.5 s, sys: 5.48 s, total: 24 s Wall time: 2.81 s # pandas (default engine) I also tried with pandas default csv parser, but I got an OOM after 20 seconds or so. I have 16GB of RAM. ```
- henrydark 4y agosuper cool. How many cores you got? Arrow has one thread serially chunking the data, and uses the rest of the cores to process the chunks in parallel. On large machines (and files, with fast io) the chunking becomes the bottleneck. Is this the same for polars?
- ritchie46 4y agoI have got 12 cores. We partition the work of over all threads, I think we are mostly IO bound as we don't use async, but we do require some pre-work before we can start reading.