7 ms·
Papers today are longer than ever and full of jargon and symbols. They depend on chains of computer programs that generate data, and clean up data, and plot dat
by PhilipVinc 5y ago
Papers today are longer than ever and full of jargon and symbols. They depend on chains of computer programs that generate data, and clean up data, and plot data, and run statistical models on data. These programs tend to be both so sloppily written and so central to the results that it’s contributed to a replication crisis, or put another way, a failure of the paper to perform its most basic task: to report what you’ve actually discovered, clearly enough that someone else can discover it for themselves.
- rtkaratekid 5y agoI think this almost every time I read the paper. It’s like Linus’ “show me the code.” I just want papers now to “show me the data and the code.” And include a discussion about why these results are important. I think it’s a great time for the scientific community to improve transparency on these fronts. Sincerely, someone who reads a lot of research but contributes none because I’m an amateur. Edit: when I say data, I mean the raw data.
- netizen-936824 5y agoRaw data can be on the order of terabytes, not that it can't be shared but this is a real barrier when it comes to raw data
- someguydave 5y agoI guess we should stop trying because datasets are big
- jrichardshaw 5y agoThe GP is making a completely legitimate point here that broad sharing of large raw datasets is pretty hard, but I don't think anyone is arguing we should give up. Here's a few thoughts, though they're more directed at the general thread than the parent. In my case I'm currently finishing up a paper where the raw data it's derived from comes to 1.5 PB. It is not impossible to share that, but it costs time and money (which academia is rarely flush with), and even if it was easy at our end, very few groups that could reproduce it have the spare capacity to ingest that. We do plan to publicly release it, but those plans have a lot of questions. Alternatively we could try to share summary statistics (as suggested by a post above), but then we need to figure out at what level is appropriate. In our case we have a relevant summary statistic of our data that comes to about 1 TB that is now far easier to share (1 TB really isn't a problem these days, though you're not embedding it in a notebook). But a large amount of data processing was applied to produce that, and if I give you that summary I'm implicitly telling you to trust me that what we did at that stage was exactly what we said we'd done and was done correctly. Is that reproducibility? You could also argue this the other way. What we've called "raw data" is just the first thing we're able to archive, but our acquisition system that generates it is a large pile of FPGAs and GPUs running 50k lines of custom C++. Without the input voltage streams you could never reproduce exactly what it did, so do you trust that? Then you're into the realm of is our test suite correct, and does it have good enough coverage? I think we have a pretty good handle on one aspect of this, is our analysis internally reproducible? i.e. with access to the raw data can I reproduce everything you see in the paper? That's a mixture of systems (e.g. configs and git repo hashes being automatically embedded into output files), and culture (e.g. making sure no one things it's a good idea to insert some derived data into our analysis pipeline that doesn't have that description embedded; data naming and versioning). But the external reproducibility question is still challenging, and I think it's better to think about it as being more of a spectrum with some optimal point balancing practicality and how much an external person could reasonably reproduce. Probably with some weighting for how likely is it that someone will actually want to attempt a reproduction from that level. This seems like the question that could do with useful debate in the field.
- deleted 5y ago[deleted]
- someguydave 5y agowhy not purchase a sufficient number of tapes or drives to capture the data and deposit it at the university library? certainly sharing apparatus is hard but you could release the schematics, board designs and BOMs of the electronics involved. The problem now is that 1) very few even try to reproduce 2) very little money is available for reproduction fixing those incentives would help alot.
- chongli 5y agoThe whole point of the field of statistics is that you can carry out statistical tests and analysis on a sample; you don’t need all of the data.
- JBorrow 5y agoIt doesn’t really work like that. For instance, imagine you have a simulation with billions of particles in it. To construct reduced data you may need to use many fields (position, temperature, composition) of all particles over many outputs (usually at different times).
- chongli 5y agoIn that case you shouldn’t need to ship the data at all. Just include the code for the simulation and let the rescuers run it to generate the data themselves.
- JBorrow 5y agoSorry I'm a bit late to this, but those simulations take 10s - 100s of millions of Cpu hours (i.e. costs of millions - 10s of millions of dollars), so that's not practical.
- bloak 5y agoI think in astronomy they generate tens of terabytes per night and an experiment may involve automatically searching through the data for instances of something rare, like one star almost exactly behind another star, or an imminent supernova, or whatever. To test the program that does the searching you need the raw data, which until recently, at least, was stored on magnetic tape because they don't need random access to it: they read through all the archived data once per month (say) and apply all current experiments to it, so whenever you submit a new experiment you get the results back one month later. I like the idea of publishing the data with the paper but it's not feasible in every case.
- zmb_ 5y agoThere are also legal and privacy concerns. I've worked on a few research papers where exactly one researcher had access to the data under a very strict NDA. And even they did not get full access to the raw data, only the ability to run vetted code against it and some subsets for development. This is because the datasets were subscriber logs from mobile operators. They are both highly privacy sensitive and contain sensitive business knowledge. There is no way they will ever get published, even in some anonymized form. Ultimately it always comes down to trust. You need to convince your peer reviewers to trust you that you have correctly done what you have claimed to have done. Of course, even when you publish datasets, you need to convince the peer reviewers to trust you that you didn't fake the data.
- d110af5ccf 5y agoI agree that they should ideally come with raw data along with all code that was used to process it to produce the results as presented. > but contributes none because I’m an amateur I don't mean to be rude but it seems relevant to point out. Papers aren't written for the benefit of amateurs. They're written for experts who actively work in that specific field. I don't think there's anything wrong with that.
- rtkaratekid 5y agoYeah I agree, but I read papers mainly in domains I do have university level degrees in. So while I’m not as expert as a lifetime professor, I do know the fields relatively well. And I don’t think it’s rude, that’s why I included that statement!
- jiggunjer 5y agoIt's not common for ethics boards to permit patient scans to become public domain, even when anonymized.
- chrisseaton 5y ago> I just want papers now to “show me the data and the code.” But the code is secondary to the idea. The idea and the discussion around how it was arrived at and what it means is the key thing. The code is just there to implement it. You could code the same idea ten different ways.
- LudwigNagasena 5y agoThe code is there to show that the idea is worth the discussion around.
- chrisseaton 5y agoWould the idea be valuable without the code? Yes.
- vlunkr 5y agoNot if the code is wrong, and therefore the conclusion may be wrong. I'm no scientist, but I don't think the point of scientific papers is to get unfounded ideas out into the world.
- chrisseaton 5y agoI can list many major influential papers in computer science that described an idea and didn't really give any concrete code, where we're still using the idea today. For example the paper on polymorphic inline caching, which is the key idea for the performance of many programming languages today, just described the idea, and didn't present any code. How was it evaluated? People sat and thought about it. Holds up today. You can reason about an idea through other things than concrete code. Code is transient and incidental. Ideas persist.
- d110af5ccf 5y agoI think you're talking past each other. Both are true under different circumstances. In some cases an abstract idea is the important takeaway. In other cases the central point of a paper is to present conclusions that were arrived at based on analysis of some dataset. If the code used to generate or analyze the dataset is wrong then conclusions based on it likely worthless.
- 14 5y agoI remember in grade school reports seemed to logical. Propose something, create a hypothesis what you think you will see, record your data and what you observe during the experiment, summarize the results as to what actually happened vs what you initially expected. Most papers now seem like a foreign language and I can only glimpse at what is happening relying on some math genius to reply the significance.
- haihaibye 5y agoOne of the first suggestions I have is to use source control and store the Git hash of code used to generate data. A few times I've heard back "we don't have time for that" - pretty easy to see how the replication crisis flows from processes like that.
- jiggunjer 5y agoEven then, the entire software environment and even the compiler choice or different hardware could cause numerical differences.
- _Microft 5y agoPapers aren’t pop-sci articles, they do not target an audience that does not knows anything about the field yet. They are from experts for experts. If someone wants to familiarize themselves with the language, symbols and methods of a field, a textbook is a better thing to start with. Over time they will also learn the shared knowledge of the field that isn’t even mentioned in these articles.
- ketozhang 5y agoCertain scientific software packages (e.g., Tensorflow, pymc3, etc) do have frameworks that you follow to return pipeline and result objects that follow some data model that others can learn quickly (e.g., an arviz::InferenceData result object). I wish there was a more extensive framework where this is applied end-to-end from data input, to library components in a pipeline processes, to the result, and then to plot.
- fho 5y agoWorse in that a lot of researchers actually have only the slightest grasp on statistics. To the point that I would assume that a lot (1 in 20? :-)) papers will contain an error in their statistical analysis of their results.