6 ms·
The benchmarks[1] from polars' websites seems really promising, and I am going to try it on my next jupyter notebook. But what I wonder really is about parsing
by chazeon 4y ago
The benchmarks[1] from polars' websites seems really promising, and I am going to try it on my next jupyter notebook. But what I wonder really is about parsing time, this is not shown in bench marks.
I my previous project, I use pandas' read_table to avoid for loop to parse molecular dynamics trajectories LAMMPS. The trajectories are pure text file with size of a few Gigs. Compared to a pure Python implementation ~3-5 min, pandas reduce it to 20 secs. I wonder how Polars works.
To be honest, compared to parsing, data manipulation such as the benchmark showcased, is just not so costly.
[1]: https://www.pola.rs/benchmarks.html https://www.pola.rs/benchmarks.html
- maegul 4y agoMy understanding is that polars can get fast(er) parsing mainly/only by leveraging arrow and parquet files.
- Jweb_Guru 4y agoAs far as I know, it's much faster than Pandas (and other competitors) at parsing plain CSVs as well.
- ritchie46 4y agoAuthor of polars here. That's incorrect. I have put a lot of effort in polars' csv parser for instance and it is one of the fastest csv parsers out there. This has nothing to do with leveraging arrow. We control all our parsers and read directly into arrow memory. This differs from pandas, which utilizes pyarrow for reading parquet and then finally has to copy the arrow memory over to pandas memory.
- maegul 4y agoApologies! I got that from some blog post somewhere I believe, not from any personal wackereren our judgement (which I should have signalled better in my post). Nonetheless, would reading a parquet file with polars be faster than reading a csv? Also thanks for polars! Great contribution data science!
- ritchie46 4y ago> Apologies! I got that from some blog post somewhere I believe, not from any personal wackereren our judgement (which I should have signalled better in my post). No worries. :) Most blogs on the topic I encounter in the wild make incorrect claims, I understand the confusion. > Nonetheless, would reading a parquet file with polars be faster than reading a csv? Yes, much faster. Please don't use the csv format for anything of a reasonable size. It is a terrible format to process and very ambiguous.
- ritchie46 4y agoAuthor of polars here. If we would have included csv parsing in the benchmark, polars would have come out much better than it already has.
- henrydark 4y agoBetter than the apache arrow c++ implementation that's in pyarrow?
- ritchie46 4y agoIt is on par. Just did a local benchmark on the new york taxi dataset. ``` # polars We should not rechunk as that is not what pyarrow does. Rechunking is a different operation and is also optimized away in many lazy operations. %%time pl.read_csv("csv-benchmark/yellow_tripdata_2010-01.csv", rechunk=False, ignore_errors=True) CPU times: user 17.4 s, sys: 2.42 s, total: 19.8 s Wall time: 1.89 s %%time pa.csv.read_csv("csv-benchmark/yellow_tripdata_2010-01.csv") CPU times: user 17.1 s, sys: 3.07 s, total: 20.1 s Wall time: 1.99 s # pyarrow embedded new lines (e.g. valid csv files) Polars by default allows embedded new line characters. Pyarrow does not, if we tell it to do so, polars is faster on my laptop. %%time pa.csv.read_csv("csv-benchmark/yellow_tripdata_2010-01.csv", parse_options=pa.csv.ParseOptions(newlines_in_values=True)) CPU times: user 17.8 s, sys: 3.09 s, total: 20.9 s Wall time: 2.72 s # pandas (pyarrow engine) And pandas itself has to pay for the copy from pyarrow to pandas: %%time pd.read_csv("csv-benchmark/yellow_tripdata_2010-01.csv", engine="pyarrow") CPU times: user 18.5 s, sys: 5.48 s, total: 24 s Wall time: 2.81 s # pandas (default engine) I also tried with pandas default csv parser, but I got an OOM after 20 seconds or so. I have 16GB of RAM. ```
- henrydark 4y agosuper cool. How many cores you got? Arrow has one thread serially chunking the data, and uses the rest of the cores to process the chunks in parallel. On large machines (and files, with fast io) the chunking becomes the bottleneck. Is this the same for polars?
- nicoburns 4y agoYou only need to avoid for loops because Python is so slow. The fantastic thing about Rust is that for the most part you can just use a for loop and it’ll be fast by default (there are still tricks to speed things up that you may need in some scenarios)
- agoose77 4y agoThis is certainly a motivation, but there are other reasons to avoid loops. In many domains, array-at-a-time abstractions offer a higher-level view of the problem. That's why something like `xtensor` exists for a NumPy-like API in C++.