3 ms·
I recently reached the limits of Pandas running on my 2020 16gb M1. Counting the number of times an element appears in a 1.7B row DataFrame using `df.groupby().
by recursive4 3y ago
I recently reached the limits of Pandas running on my 2020 16gb M1. Counting the number of times an element appears in a 1.7B row DataFrame using `df.groupby().size()` would consistently exceed available memory.
Rust Polars is able to handle this using Lazy DataFrames / Streaming without issue.
- sweezyjeezy 3y agoFWIW I think df.column.value_counts() is better to use here in pandas.
- recursive4 3y agoIt unfortunately also exceeded available memory. A basic approach which worked was sequentially loading each df from the filesystem, iterating through record hashes, and incrementing a counter; however the runtime was an order of magnitude greater than my final implementation in Polars.