4 ms·
I agree with your critiques of Python... Could you please post some example of code/operations which are very natural in R but unnatural in Python/Pandas? I'm c
by ced 11y ago
I agree with your critiques of Python... Could you please post some example of code/operations which are very natural in R but unnatural in Python/Pandas? I'm curious to see what I'm missing out on.
- vegabook 11y agoWell, I use both, and I can do everything in Python that I can do in R. However here are some things which will give you a flavour of R's more consistent, data-first nature: > rollapply(some1000x10matrix, 200, function(x) eigen(cov(x))$values[1], by.column = FALSE) # get the first eigenvalue rolling 200x10 window. >>> # impossible in Python unless using ultra-complex Numpy stride tricks. > dim(someMatrix) >>> someMatrix.shape > head(someMatrix) >>> someMatrix.head() # notice consistent function application in R, whereas in Python, mixed attribute / function? So we're on OO land and I must know if it's an attribute or a function.... > rollapply(some1000x2matrix, 200, function(x) {linmod <- lm(x[, 1] ~ x[, 2]); last(linmod$residuals) / sd(linmod$residuals)}, by.column = FALSE) # get the z score in one multi-step function. >>> Impossible in python without For loop as lambdas cannot be multi-statement. > native indexing using [] brackets by index number, or index value, or boolean. All vectors. >>> pandas loc/iloc/ix mess. > ordered lists (python dict) by default, so boolean or index subsection easy even when data is hierarchical, not tabular >>> easy bugs due to unordered nature of dicts; must import some different module and then still can't vector index it. It's all summed up by this: > c(1, 2, 3) * 3 [1] 3 6 9 >>> [1, 2, 3] * 3 [1, 2, 3, 1, 2, 3, 1, 2, 3] # wrong! Need rescuing by Numpy! And then there's CRAN. Just last night someone told me about "nowcasting" which uses "MIDAS regression". A relatively new technique. Google it for R (full package available), Google it for Python (Matlab comes up ;-). And I'm not even going to start on graphics. Seaborn and bokeh are valiant efforts, but they're still 80% of what ggplot and base graphics can do, especially, at the multidimensional scale. That last 20% is often all the difference between meh and wow. That said, I do appreciate Matplotlib's autos rescaling of axes when adding data. Python charts aren't as pretty nor capable of complexity (for similar effort), but they're arguably more dynamic. Now don't get me wrong. The converse list for Python would be much longer, because it's more general purpose, and it kills R outside of data science. I wrote 10k loc in R for a semi-production and it was horrible because it does not have the CS tools for managing code complexity, and it really is slow at certain things. R is more focused on iterative, exploratory data science, where it excels.