3 ms·
I think it really depends on the scale of data. If you're dealing with anything less than a GB, it probably doesn't matter all that much, but once you're dealin
by ajoseps 4y ago
I think it really depends on the scale of data. If you're dealing with anything less than a GB, it probably doesn't matter all that much, but once you're dealing with larger datasets there is a pretty massive difference with using vectorized operation. Some of the pandas dataframes methods map to underlying numpy ones, but I don't believe that is always the case
- oneoff786 4y agoWith the availability of things like pyspark the grey zone between pandas scale and pyspark scale is small and uncommon though. Especially for the awkward tabular data manipulation tasks where you actually need to be mapping custom functions and what not. Pandas built ins can cover everything with good performance except dataframe.apply imo That and using dicts as maps.