3 ms·
I also use R for any heavy data manipulation, but I primarily use the data.table package. The efficiency that both of these packages unlock is absolutely unpara
by CapmCrackaWaka 5y ago
I also use R for any heavy data manipulation, but I primarily use the data.table package. The efficiency that both of these packages unlock is absolutely unparalleled in any other tabular data manipulation library, in any other language that I have used. And R has the top 2!!
My skin writhes every time I need to type:
table.loc[(table.column > 2) | (table.column2 < 3)].reset_index(drop=True)
when I want to subset a table.
- alexilliamson 5y agoNot to mention the auto complete that comes with RStudio. Is there any way to get equivalent functionality in Jupyter?
- CapmCrackaWaka 5y agoI use pycharm which has decent autocomplete. Pycharm has its own issue though, it fills out its autocomplete info by looking at the function that created the object, not the object itself. So if a function can return different types, autocomplete won’t work. That’s caused me quite a bit of pain.
- discordance 5y agoIf you set up use the Jupyter extension[0] and open your notebooks in VS Code you get Intellisense (code completion, method info and hints etc). 0: https://marketplace.visualstudio.com/items?itemName=ms-toolsai.jupyter https://marketplace.visualstudio.com/items?itemName=ms-tools...
- claytonjy 5y agoIME this is strictly worse than the RStudio experience; most of the time I hit tab in a VSCode notebook I get way too many options that IMO are clearly not what I want, though at some level it's more pandas' fault (too many methods & attributes even before attaching every column name as an attribute) than VSCode or Intellisense.
- ogogmad 5y agoIn addition to the sibling answers: If you know the class of an object, let's say Class, but the object has yet to be "constructed" so that IPython can correctly infer its type, you can type `Class.[TAB]` in IPython and look at its methods. For example, in Sympy, you have a matrix type called Matrix. You can do `(A * B).diagonalize()`, or alternatively you can do `Matrix.diagonalize(A * B)`, which has some advantages because doing `(A * B).[TAB]` does nothing useful because Python can't infer types. You can also do the same for modules. `ModuleName.[TAB]` To be honest though, I found the experience smoother in R for some reason.
- sntscy 5y agoCan't get around resetting the index, as far as I know, but for the filtering you can also do, `table.query("column > 2 and column2 < 3")`
- claytonjy 5y agodo any IDEs help you autocomplete the string argument? that's a bit part of the non-standard evaluation magic in the tidyverse these days; you can use bare, _unquoted_ names, _and_ get excellent autocomplete, at least in RStudio.
- _dain_ 5y ago.loc lets you supply a callable, so you can write: table.loc[lambda df: df["column"].between(2, 3, inclusive="neither")] this is useful when your dataframe has a long name, or when you have some long method chain and you need to subset at the end: table.foo().bar().baz().loc[lambda df: ...] it is still more verbose, but I actually prefer always providing column names as strings. it's more explicit. I don't like R's environment-manipulation metaprogramming magic where you can give column names as symbols. as for resetting the index all the time, this is something of an antipattern. if you set up your index right beforehand it isn't necessary so often.