6 ms·
While I absolutely agree that scientists need more formal training in data visualization, the claim that "few scientists take the same amount of care with visua
by rflrob 7y ago
While I absolutely agree that scientists need more formal training in data visualization, the claim that "few scientists take the same amount of care with visuals as they do with generating data or writing about it" doesn't ring true to me about the scientists in my field (genetics/genomics). There is a widespread recognition that effective figures are the strongest way to communicate your message, the question is what is the the best way to create those effective figures. Part of the problem is that as datasets get bigger, it's rarely sustainable to put a lot of care into each and every version of a plot, but automating creation of figures is really hard.
If I were going to start designing a course in creating scientific figures, I think I'd have a roughly even split between the psychophysics of visual perception (e.g. distinguishing between similar quantities of lengths/angles/colors/etc; designing for color-blind readers; ) and hands-on work in a real programming environment turning data to figures.
- puttermesser 7y agoRelatedly, here’s some cool visualization work that comes from genetics data viz. http://scalable-insets.lekschas.de/ http://scalable-insets.lekschas.de/
- Thriptic 7y agoI agree with everything you said and can confirm that in our (translational vascular bio) lab the bulk of effort spent on paper drafting was concerned with creation of high quality figures. Many scientists (myself included), will read a paper abstract and then head straight for the figures as they usually contain the highest density of data for the reader.
- psalminen 7y agoThe course you are describing is one which I took. It was called "Scientific Visualization", and was a CS course mainly taken by science majors. The meta-information describing color schemes and scales was by far the most interesting part of it.
- jfim 7y ago> Part of the problem is that as datasets get bigger, it's rarely sustainable to put a lot of care into each and every version of a plot, but automating creation of figures is really hard. That's actually a good reason to learn R and the ggplot2 package. Whenever I write a paper, what I do is that I make a quick shell script that invokes Rscript, with a simple R program that takes a CSV file and outputs a PDF file of the plot, which can be automatically loaded in LaTeX. Whenever the data changes, it's just a matter of updating the CSV file and running the script that rebuilds the figures and the LaTeX document. As an added bonus, it makes keeping the data with the paper easy, since they're part of the same source control repository.
- airstrike 7y agoIs there a reason you don't use RMarkdown in RStudio? It's built precisely for this use case
- jfim 7y agoMostly because I wrote those scripts many years ago and I've been reusing them since. I'll look into RMarkdown for the next paper, thanks!
- samch93 7y agoI can totally recommend knitr by the amazing yihui xie (https://yihui.org/knitr/ https://yihui.org/knitr/ ) for this use case. It allows you to write R code chunks directly in your latex code and the output of the code (tables, plots) is then directly inserted in your pdf after compiling. Together with git and docker this gives a fully reproducible workflow!
- noobermin 7y agoI'd say there is a range, as there always is. I've seen fantastic and clear visuals of FDTD simulations of laser interactions in my field in 3d, then, I've seen jet used to pcolor data that ranges from positive to negative. I would say though that a good enough fraction (not sure if it's greater than 50% but I wouldn't be surprised if it was) do care about making good figures.