3 ms·
Series.map is vectorized. Pretty much everything you need in pandas is as performant as you ought to need for doing tabular data manipulation in Python. Except
by oneoff786 4y ago
Series.map is vectorized.
Pretty much everything you need in pandas is as performant as you ought to need for doing tabular data manipulation in Python. Except dataframe.apply
- _Wintermute 4y agoIt is not. df = pd.DataFrame({"foo": np.random.randn(100000)}) pandas map: df["foo"].map(lambda x: x * 2) 18.1 ms ± 109 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) pandas apply: df["foo"].apply(lambda x: x * 2) 17.9 ms ± 46.6 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) Vectorised function, using underlying numpy operations: df["foo"] * 2 267 µs ± 11.8 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)
- lmeyerov 4y agoUse numba and it is, including on GPU :)
- lcvriend 4y agoIf by "vectorized" you mean: "able to delegate the task of performing mathematical operations on the array's contents to optimized, compiled C code." then I do not think you are correct (unless perhaps you are supplying map with a dict or Series). Series.map is not compiling your lambda's to C and running it. If there is a built-in method available it usually will be faster. Notable exception are pandas str methods which devolve into Python code but generally with more overhead than map/apply.
- oneoff786 4y agoI mean not writing your own loops. A built in function is indeed better. But usually not what you need. And also, readability > speed when the execution time is trivial, which is it probably should be for pandas scale data.
- fumeux_fume 4y agoHey, I also use Pandas every day and I would definitely recommend keeping up with your Numpy skills sharp. A lot of Pandas is built on top of Numpy so there's one good reason, but another is that it would prevent you from footguns like thinking Series.map is vectorized.