5 ms·
I think this is a great idea theoretically, but in reality for most papers I don't want to see the data/underlying code. While it would be great to publish data
by sam-2727 5y ago
I think this is a great idea theoretically, but in reality for most papers I don't want to see the data/underlying code. While it would be great to publish data/code with the paper (in the field I've worked on the most, astronomy, most data is already published with the paper anyways), I don't want/need to look through a notebook with the underlying code of the paper in order to just read the intro/conclusions (and maybe one key methods section). Interactive figures are a great idea, but again, oftentimes I don't really care to interact with the figure, or fiddle stuff around, I just want to know why the paper is important and how I should use its conclusions. The two-column format of most papers is very useful for skimming. So instead I would argue notebooks shouldn't replace papers, but supplement them (as they sometimes do already, in fact, but perhaps journals could make it an actual requirement to create a supplementary notebook).
As the article mentions, scientific fields are gigantic nowadays, and skimming papers is critical when you're citing 100+ references in your paper.
- deleted 5y ago[deleted]
- TuringTest 5y agoThe point of interactive notebooks is not seeing and having access to all the data - it's seeing the abstractions at work, having a direct grasp of how they act on particular examples as an aid to understand their formal definition. Nothing prevents you from having two-column notebooks, if you find that advantageous, as well as abstract and conclusions sections. The part that you don't get with static paper is that of navigating the abstraction ladder[1] up and down with direct manipulation aids, instead of having to work it all in your head or by following dense detailed paragraphs. [1] As also explained by Bret Victor in http://worrydream.com/LadderOfAbstraction/ http://worrydream.com/LadderOfAbstraction/
- GracefullyBlind 5y agoI think having the ability to focus on the things you care about the paper mostly is what would be more beneficial for all readers. You care more about an overview? You can easily find it (perhaps with graphics and walkthroughs), you care more about proofs? Then you can get them, what about code and experiments? And so on and so forth. Readability and scalability is about making all this data available in the publication record, but easy to navigate for whoever is looking for whatever.
- nextos 5y agoEMBL-EBI and others had some RDF-related effort to provide machine readable abstracts, which I thought was a really cool idea. IMHO, the biggest problem with papers is politics and reviews. In many top journals like Nature there's no double-blind review (actually in Nature it's now optional but big groups never use it). And even if there was double-blind review, referees have no skin in the game. So the usual outcome is to get reviewed by a big name in your field, who is actually interested in controlling research trends and killing "competitors". This is hindering progress and hurting new ideas. For example, proponents of Alzheimer's disease being caused by an infection or dysbiosis have had a hard time to do research, get grants and publish articles during the last 2 decades. Despite their theory is able to explain the etiology quite well, unlike competing alternatives. Another problem is that to publish in good journals you need cool results. Cool results are rare, but Nature, Science, Cell et al. are full of articles every month. So, most groups are overselling and misreporting things. Research fraud, p-value hacking and data manipulation are really common.
- ansgri 5y agoIt’s not really possible to conduct double-blind reviews in most cases: authors or at least the group can often be easily guessed from the list of references, “in our previous work…”, and research domain and approach in general.
- jamessb 5y agoSensible anonymisation policies prevent people from referring to "our previous work" in submissions - e.g., the policy for CHI [1] states: > We do expect that authors leave citations to their previous work unanonymized so that reviewers can ensure that all previous research has been taken into account by the authors. However, authors are required to cite their own work in the third person, e.g., avoid “As described in our previous work [10], … ” and use instead “As described by [10], …” However, it is true that things like choice of research questions, approach, and equipment used can be quite suggestive of the authors' identity. [1]: https://chi2020.acm.org/authors/papers/chi-anonymisation-policy/ https://chi2020.acm.org/authors/papers/chi-anonymisation-pol...
- 5y ago
- bloaf 5y agoYou should always want to have the underlying code available. Without the exact procedures they used to process their data, the only kind of "using their conclusions" you can do is the superficial "take it at face value" kind. So many important details get hand-waved away in papers that say things like "we used the well known blahblahblah method to analyze the data." If you do it right, the code should in no way interfere with your ability to read abstracts.
- deleted 5y ago[deleted]
- eesmith 5y agoI think I can convince you otherwise. If I publish a paper saying I have an algorithm which can factor large composites, and in the paper publish the factors to all of the RSA numbers listed at https://en.wikipedia.org/wiki/RSA_Factoring_Challenge https://en.wikipedia.org/wiki/RSA_Factoring_Challenge , then I think people will take it seriously, and not consider it at the superficial level. Even if I don't publish the algorithm. ("Because of the security implications of this work, I have decided to withhold publication for a year.") Furthermore, some things are worth publishing even if the methods was "it came to me in a dream" à la Kekulé's snake. If you can demonstrate a sorting network of size 47 for n=14 input (which is the known lowest bound) then you can publish that exemplar, even without publishing the method used to generate it. (If you used computer assistance then that method would likely also be publishable, but that's a different point. Newton famously used the calculus to solve problems, but published their proofs using more traditional approaches.) If you can come up with a protein model that is a significantly better fit to the X-ray diffraction data, then that's publishable too, no matter how you came up with that model. In all of these cases, there are ways to verify the validity of the results without reproducing the methods used to come up with the result.
- deleted 5y ago[deleted]
- lonesword 5y agoThis won’t work for empirical research. I vividly recall weeks spent trying to reproduce a paper on information retrieval (a deep learning model). What saved me is skimming through the author’s codebase and chancing upon an undocumented sampling step. They were only using the first and last passage in a document as training data and uniformly sampling from 10% of the remaining passages, and the paper didn't mention this. I adopted their sampling strategy, and i was able to obtain their results. My argument is that there are nuances and subtleties that are often omitted in a paper (accidentally or otherwise), but are nevertheless required to reproduce the research.
- Helmut10001 5y agoI have made an experiment with my last paper: Write everything from scatch in Jupyter Notebook, including data preprocessing and generation of all figures (etc.) (10 Notebooks in total). Start of the conceptualization was in 2017, we just submitted it 2 weeks ago (it got desk rejected for not fitting the journals topic). I learned a lot and it was definitly worth it. The next paper will be easier with this knowledge. Nonetheless, there is an overhead and I feel that this overhead is not valued with the current makeup of journals, where you really need to dig deep to find any supplementary materials.
- rscho 5y agoDepending on the stuff you do, emacs org-mode is worth a shot. I write all my papers in org.
- fho 5y agoI did [something similar] too when I started my PhD ... I had one Makefile managed project that ran everything with dependencies. From raw data, to figures and even embedding the numbers into the final, Latex-based PDF. My supervisor manually copied all of the text from my PDF into a word document on his first revision ...
- RandomLensman 5y agoMoving the burden of assessing everything in a paper completely to the reader is an interesting idea but seems somewhat a step back when at the same time good and curated data gets ever more expensive. So the market for validated results is already not bad where those results "matter". And not every paper has a lot of code or data associated with it. If you do experiments on organisms etc. then there is so much happening in the actual lab work - where would that go? Endless hours of video documentation?
- atoav 5y agoIn a ipython notebook you can fold away "blocks" of code, that means you can have everything there that produces the graphs and still be able to look under the hood if you like to. Isn't that the practical part about digital technology? That you are not limited to one view?
- Apofis 5y agoI don't know about this... considering how well some of you guys write papers, I'd rather look at the results and the code than read your paper.