4 ms·
I'm not totally sure whether this analysis captures the true extent which R vs SAS vs SPSS is used. If I use R for a plot, or a simple bit of regression, or an
by Malarkey73 10y ago
I'm not totally sure whether this analysis captures the true extent which R vs SAS vs SPSS is used.
If I use R for a plot, or a simple bit of regression, or anova, or even cross-validation. I don't reference it in a paper. I only cite it if there is a package designed for a particular type of data (e.g. a Bioconductor package) or something a bit more esoteric (e.g. apcluster). About 95% of the work is data munging and - sorry Hadley - I don't cite dplyr, purrr, magrittr etc...
However I have notice that in clinical trial or small social science papers simple analyses of this type are often cited as being done in SPSS or SAS. I think this just reflects the fact that non specialist data analysts are more likely to cite SAS or SPSS for simple procedures such as graphs or anova as an appeal to authority.
So I reckon the data may reflect a trend but tells us little about the true levels.
- ececconi 10y agoThis is especially true in undergrad papers.
- jsprogrammer 10y agoShouldn't your full source be available; which would explicitly record your dependencies?
- griffinmichl 10y agoSource code really should be available, but it almost never is. The peer review process is, in my opinion, quite flawed. While your paper's high level content gets reviewed, no one actually looks at your code and data to ensure that you didn't forget to carry the one. Your analysis could be totally wrong, but reviewers only review what you say you did, not what you actually did.
- chrisamiller 10y agoAny paper I peer review better have source code available or they will hear about it. That said, yes, there isn't time (or funding) to actually re-run the entire analysis.
- hooloovoo_zoo 10y agoWhy don't you cite the packages you use?
- noelsusman 10y agoNobody cites every package they use, it's not feasible. I use a lot of packages, and some journals have a limit on the number of citations you can have. I only cite packages when it provides specialized statistical functionality. For example, I do a lot of work with data from complex surveys, and I always cite Lumley's survey package because without it I wouldn't be able to do the work. On the flip side, I use Hadley's readr package extensively because I think his I/O functions are more sane than the defaults. I'm not going to cite readr in every paper I write just because I'm too lazy to type stringsAsFactors = FALSE when I read a csv file.
- hadley 10y agoTo me, citing readr feels like citing the company that made your pipettes. It's just useful infrastructure.
- chrisamiller 10y agoOur genomics workflows use dozens of packages even before I get the data and start really doing analysis, statistics, and plots. It's just not feasible to cite every bit of code we use (though we certainly point people to the higher level routines, which they can use to see what was run/how to reproduce the results).
- _of 10y agoI don't cite packages like plyr and ggplot, but I do cite what program I use for statistics (R, etc) and its version.