3 ms·
I just did one using fletcher (https://github.com/xhochy/fletcher https://github.com/xhochy/fletcher), a library that extends pandas arrays with arrow arrays:
by batxu 8y ago
I just did one using fletcher (https://github.com/xhochy/fletcher https://github.com/xhochy/fletcher), a library that extends pandas arrays with arrow arrays:
n = 2**25
data = np.random.choice(list(string.ascii_letters), n)
df = pd.DataFrame({
'string': data,
'arrow': fl.FletcherArray(data),
'categorical': pd.Categorical(data),
'ints': np.arange(n)
})
for a groupby operation on string/arrow/categorical and sum the ints the results are:
- String: 1.58 s
- Arrow: 886 ms
- Categorical: 406 ms
The base type of this string FletcherArray array is <pyarrow.lib.ChunkedArray, so maybe there is another more convenient arrow array type or the bottleneck is in upper layer around the groupby operation (since it only is twice as fast). Any way, for this kind of operation the Categorical is the winner and double speed is quite a notable improvement.
- gulda 8y agoNice!