3 ms·
If you want to read csvs fast in Python, you should consider using the Apache Arrow[0] csv reader. Depending on your number of CPU cores it can be 10x-20x as f
by RobinL 6y ago
If you want to read csvs fast in Python, you should consider using the Apache Arrow[0] csv reader. Depending on your number of CPU cores it can be 10x-20x as fast as the native pandas reader. [1]
More broadly, because Arrow is cross platform it can give you similar performance in many languages. And once the dataframe is in memory, you can share it between languages with no need for serialisation and deserialisation.
[0] https://arrow.apache.org/ https://arrow.apache.org/
[1] https://youtu.be/fyj4FyH3XdU?t=1036 https://youtu.be/fyj4FyH3XdU?t=1036
- karbarcca 6y agoI agree; if I needed to parse CSVs in python and could utilize the arrow format, I would definitely use pyarrow. I actually recently finished support for reading/writing the arrow format in Julia (https://github.com/JuliaData/Arrow.jl https://github.com/JuliaData/Arrow.jl), and it's automatically integrated with the CSV.jl package; so you can do `Arrow.write("data.arrow", CSV.File("data.csv"))` and convert a csv file to arrow format directly. I'm very bullish on arrow as a standard binary data format for the future.
- boxed 6y agoPandas is pretty slow and since it loads into memory it can be totally infeasible for even relatively small data sets. The csv module it what one should compare it to imo.
- RobinL 6y agoApache Arrow reads csvs into memory in Arrow format, not pandas format. They are independent libraries that do different things. In Arrow, it is possible to read a csv in batches, obviating memory problems. See 'incremental reading' - http://arrow.apache.org/docs/python/csv.html http://arrow.apache.org/docs/python/csv.html