4 ms·
Author of polars here. That's incorrect. I have put a lot of effort in polars' csv parser for instance and it is one of the fastest csv parsers out there. This
by ritchie46 4y ago
Author of polars here. That's incorrect. I have put a lot of effort in polars' csv parser for instance and it is one of the fastest csv parsers out there. This has nothing to do with leveraging arrow.
We control all our parsers and read directly into arrow memory. This differs from pandas, which utilizes pyarrow for reading parquet and then finally has to copy the arrow memory over to pandas memory.
- maegul 4y agoApologies! I got that from some blog post somewhere I believe, not from any personal wackereren our judgement (which I should have signalled better in my post). Nonetheless, would reading a parquet file with polars be faster than reading a csv? Also thanks for polars! Great contribution data science!
- ritchie46 4y ago> Apologies! I got that from some blog post somewhere I believe, not from any personal wackereren our judgement (which I should have signalled better in my post). No worries. :) Most blogs on the topic I encounter in the wild make incorrect claims, I understand the confusion. > Nonetheless, would reading a parquet file with polars be faster than reading a csv? Yes, much faster. Please don't use the csv format for anything of a reasonable size. It is a terrible format to process and very ambiguous.