3 ms·
I don't understand the advantage of the assign-method in the article you have linked to. If I would like to alter data in a column of a large dataframe, why sho
by pvitz 5y ago
I don't understand the advantage of the assign-method in the article you have linked to. If I would like to alter data in a column of a large dataframe, why should I use the assign-method (and copy data) instead of the mentioned loc-method and just overwrite the data (without a copy)? I am assuming, of course, that I will not need the original dataframe anymore.
- lmeyerov 5y agoYep, we regularly use assign & pipe to avoid the 80% case I think the article is about: ``` df2 = df.assign(new_col_1=f(df)) ``` or ``` def add_new_col_1(df): return df.assign(new_col_1=...) df2 = df.pipe(add_new_col_1) ``` We do use reified compute DAGs in some places, but by that point, we're using dask anyways, and that kind of code ends up more annoying than if we could avoid. Speaking as someone who has done years of FRP/streams/etc., if users can stick with direct control flow or things that look like it (async/await), help them do it :) RE:Immutability, it's a convention that eliminates some classes of bugs, so quite nice when a team does it. Code written by senior / reviewed teams are nice b/c these things add up to either a pleasant experience or perpetual paranoia for the day-to-day: * A lot of our code is in notebook envs, where the ability to move cells up/down matters, so we try to only do single assignments to avoid non-reproducibility bugs * Similar in production, but more about when someone comes back and wants to edit or log. having to undo the name reuse is annoying, and may lead to bugs when you don't notice it