6 ms·
R especially dplyr/tidyverse is so underrated. Working in ML engineering, I see a lot of my coworkers suffering through pandas (or occasionally polars or even b
by cye131 1y ago
R especially dplyr/tidyverse is so underrated. Working in ML engineering, I see a lot of my coworkers suffering through pandas (or occasionally polars or even base Python without dataframes) to do basic analytics or debugging, it takes eons and gets complex so quickly that only the most rudimentary checks get done. Anyone working in data-adjacent engineering work would benefit from R/dplyr in their toolkit.
- kasperset 1y agoI love R and dplyr. It is very readable and easy to explain to non-programmers. I use it almost everyday. Not exactly on the topic,I am having difficulties debugging it. May be I need to brush up on debugging R. Not sure if there is a easy way to add breakpoint when using vscode.
- JackeJR 1y agobrowser() ?
- disgruntledphd2 1y agotrace subsumes browser, it's much more flexible and can be applied to library code without editing it.
- tylermw 1y agotrace is great for shimming in your own code to an existing function, but it’s not an interactive debugging tool.
- disgruntledphd2 1y agoIt sure is. If you set the second argument to browser you can step through any function.
- wdkrnls 1y agoIs there a way to trace an attribute to a function? I couldn't find one, but curious if it exists. I seemed blocked by the fact that trace seemed to expect a name as a character string. Some functions in base R have functions in their attributes which modify their behavior (e.g. selfStart). I ended up just copying the whole code locally and then naming it, but for a better interactive experience I really wish there was a way to pass a function object as I can with debug.
- itsmevictor 1y agoHave you checked this extension? https://marketplace.visualstudio.com/items?itemName=RDebugger.r-debugger https://marketplace.visualstudio.com/items?itemName=RDebugge...
- wwweston 1y agowhat’s the story integrating R code into larger software systems (say, a saas product)? I’m sure part of Python’s success is sheer mindshare momentum from being a common computing denominator, but I’d guess the integration story is part of the margins. Your back end may well already be in python or have interop, reducing stack investment and systems tax.
- dajtxx 1y agoI am working on a system at present where the data scientist has done the calculations in an R script. We agreed upon an input data.frame and an output csv as our 'interface'. I added the SQL query to the top of the R script to generate the input data.frame and my Python code reads the output CSV to do subsequent processing and storage into Django models. I use a subprocess running Rscript to run the script. It's not elegant but it is simple. This part of the system only has to run daily so efficiency isn't a big deal.
- shoemakersteve 1y agoAny reason you're using CSV instead of parquet?
- epistasis 1y agoCSV seems to be a natural and easy fit. What advantage could parquet bring that would outweigh the disadvantage of adding two new dependencies? (One in Python and one in R)
- pjacotg 1y agoNot the op, but I started using parquet instead of CSV because the types of the columns are preserved. At one point I was caching data to CSV but when you load the CSV again the types of certain columns like datetimes had to be set again. I guess you'll need to decide whether this is a big enough issue to warrant the new dependencies.
- pletnes 1y ago
- joshdavham 1y agoTotally agreed that R is underrated. I'm sad that I stopped using it after graduation.
- vishnugupta 1y agoAs someone who is learning probability and statistics for recreation, I wholeheartedly agree. I wish I had come across R and dplyr/tidyverse/ggplot2 back in college while learning probability and stats. They were quite boring and drudgery to study because I wasn't aware of R to play around with data. Well, better late than never I guess.
- gnuly 1y agoR was the first thing we had in our syllabus for (shallow)Machine Learning. the ease of doing `model <- lm(speed~dist, cars)` and then `predict(model, data.frame(dist = c(42)))` is unparalled.
- aquafox 1y agoWhy not mix R and Python in interactive analysis workflows: 1) Download positron: https://github.com/posit-dev/positron https://github.com/posit-dev/positron 2) Set up a quarto (.qmd) notebook 3) Set up R and Python code chunks in tour quarto document 4a) Use reticulate to spawn a Python session inside R and exchange objects beween both languages (https://github.com/posit-dev/positron/pull/4603 https://github.com/posit-dev/positron/pull/4603) 4b) Write a few helper functions that pass objects between R and Python by reading/writing a temporary file.
- dkga 1y agoThis is exactly what I do for the vast majority of my academic papers. It combines the power and flexibility of R for statistics, which I agree with the upstream poster is incredibly underrated (especially with tidyverse) with python.
- Annatar 1y ago[dead]
- goosedragons 1y agoOrg mode in Emacs is even better at this IMO. Only downside is that no guarantee other people use Emacs too.
- b-rodrigues 1y agoI'm writing a package called rixpress that leverages Nix to build reproducible pipelines with targets in either R or Python Here's the github to the package https://github.com/b-rodrigues/rixpress/tree/master https://github.com/b-rodrigues/rixpress/tree/master and here's an example pipeline https://github.com/b-rodrigues/rixpress_demos/tree/master/python_r https://github.com/b-rodrigues/rixpress_demos/tree/master/py...
- p00dles 1y agoIs this what tools like Nextflow or Snakemake aim to do? I don't know, and I'm genuinely curious, because I'm starting to work in bioinformatics and doing different parts of an analysis pipeline in R and Python seems common, and, necessary really if you want to use certain packages. I'm wondering if I should devote time to learning Nextflow/Snakemake, or whether the solution that you outlined is "sufficient" (I say "sufficient" in quotes because of course, depends on the use case).
- fithisux 1y agoLife saver. I do not use the raw dataframe API, inconsistent and error prone.