4 ms·
I do almost all of my day job in pandas. I consider myself very good at it. My number one recommendation to new data scientists learning the ropes is to just no
by oneoff786 4y ago
I do almost all of my day job in pandas. I consider myself very good at it. My number one recommendation to new data scientists learning the ropes is to just not use NumPy almost at all. I’m not sure where people learn it but they do all of this complicated nonsense. Just map simple Python lambda funcs with pd.Series.map and that’s most of what you need. Memorize your pd.DataFrame methods.
If your code feels like it dealing with a matrix and not a table, it’s probably doing something funny.
- _Wintermute 4y agoYou lose a lot of performance not using vectorised functions. Maybe not an issue if you're only dealing with small amounts of data.
- oneoff786 4y agoSeries.map is vectorized. Pretty much everything you need in pandas is as performant as you ought to need for doing tabular data manipulation in Python. Except dataframe.apply
- _Wintermute 4y agoIt is not. df = pd.DataFrame({"foo": np.random.randn(100000)}) pandas map: df["foo"].map(lambda x: x * 2) 18.1 ms ± 109 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) pandas apply: df["foo"].apply(lambda x: x * 2) 17.9 ms ± 46.6 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) Vectorised function, using underlying numpy operations: df["foo"] * 2 267 µs ± 11.8 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)
- lmeyerov 4y agoUse numba and it is, including on GPU :)
- lcvriend 4y agoIf by "vectorized" you mean: "able to delegate the task of performing mathematical operations on the array's contents to optimized, compiled C code." then I do not think you are correct (unless perhaps you are supplying map with a dict or Series). Series.map is not compiling your lambda's to C and running it. If there is a built-in method available it usually will be faster. Notable exception are pandas str methods which devolve into Python code but generally with more overhead than map/apply.
- oneoff786 4y agoI mean not writing your own loops. A built in function is indeed better. But usually not what you need. And also, readability > speed when the execution time is trivial, which is it probably should be for pandas scale data.
- fumeux_fume 4y agoHey, I also use Pandas every day and I would definitely recommend keeping up with your Numpy skills sharp. A lot of Pandas is built on top of Numpy so there's one good reason, but another is that it would prevent you from footguns like thinking Series.map is vectorized.
- voxelghost 4y agoCheck out polars. Vectorized, choice between lazy optimization and eager.
- boppo1 4y agoWhat is your day job?
- oneoff786 4y agoData science consulting for business things. Most datasets < 1M rows
- eskaytwo 4y agoThose are small datasets. Also numpy != just operations on a matrix, but as mentioned above, proper vectorization can have profound speed differences for larger datasets.
- oneoff786 4y agoThey are indeed small. But most real world problems operate on small data. And if it’s bigger you probably ought to be using a different toolset than Python altogether. In my experience it is exceedingly rare to find a situation where you have > 1M rows of data and need to do tabular data manipulation, and it’s not coming from a managed database sort of setup where the manipulation could be better done in the database.
- ajoseps 4y agoI think it really depends on the scale of data. If you're dealing with anything less than a GB, it probably doesn't matter all that much, but once you're dealing with larger datasets there is a pretty massive difference with using vectorized operation. Some of the pandas dataframes methods map to underlying numpy ones, but I don't believe that is always the case
- oneoff786 4y agoWith the availability of things like pyspark the grey zone between pandas scale and pyspark scale is small and uncommon though. Especially for the awkward tabular data manipulation tasks where you actually need to be mapping custom functions and what not. Pandas built ins can cover everything with good performance except dataframe.apply imo That and using dicts as maps.