3 ms·
Has anyone done any perf analysis between this and previous versions?
by bpchaps 8y ago
Has anyone done any perf analysis between this and previous versions?
- batxu 8y agoI just did one using fletcher (https://github.com/xhochy/fletcher https://github.com/xhochy/fletcher), a library that extends pandas arrays with arrow arrays: n = 2**25 data = np.random.choice(list(string.ascii_letters), n) df = pd.DataFrame({ 'string': data, 'arrow': fl.FletcherArray(data), 'categorical': pd.Categorical(data), 'ints': np.arange(n) }) for a groupby operation on string/arrow/categorical and sum the ints the results are: - String: 1.58 s - Arrow: 886 ms - Categorical: 406 ms The base type of this string FletcherArray array is <pyarrow.lib.ChunkedArray, so maybe there is another more convenient arrow array type or the bottleneck is in upper layer around the groupby operation (since it only is twice as fast). Any way, for this kind of operation the Categorical is the winner and double speed is quite a notable improvement.
- gulda 8y agoNice!