15 ms·
The Scientific Paper is Obsolete (2018)
- ChrisArchitect 5y agoWhy'd you post this? Something new? Posted annually since original and plenty of discussion then: https://news.ycombinator.com/item?id=16764321 https://news.ycombinator.com/item?id=16764321
- dang 5y agoIf an article hasn't had significant attention in the last year or so, reposts are ok. This is in the FAQ: https://news.ycombinator.com/newsfaq.html https://news.ycombinator.com/newsfaq.html.
- ChrisArchitect 5y agoNot saying it's not ok, just want to know the poster's motivation for the share, especially if there's nothing new etc
- PhilipVinc 5y agoPapers today are longer than ever and full of jargon and symbols. They depend on chains of computer programs that generate data, and clean up data, and plot data, and run statistical models on data. These programs tend to be both so sloppily written and so central to the results that it’s contributed to a replication crisis, or put another way, a failure of the paper to perform its most basic task: to report what you’ve actually discovered, clearly enough that someone else can discover it for themselves.
- rtkaratekid 5y agoI think this almost every time I read the paper. It’s like Linus’ “show me the code.” I just want papers now to “show me the data and the code.” And include a discussion about why these results are important. I think it’s a great time for the scientific community to improve transparency on these fronts. Sincerely, someone who reads a lot of research but contributes none because I’m an amateur. Edit: when I say data, I mean the raw data.
- netizen-936824 5y agoRaw data can be on the order of terabytes, not that it can't be shared but this is a real barrier when it comes to raw data
- someguydave 5y agoI guess we should stop trying because datasets are big
- jrichardshaw 5y agoThe GP is making a completely legitimate point here that broad sharing of large raw datasets is pretty hard, but I don't think anyone is arguing we should give up. Here's a few thoughts, though they're more directed at the general thread than the parent. In my case I'm currently finishing up a paper where the raw data it's derived from comes to 1.5 PB. It is not impossible to share that, but it costs time and money (which academia is rarely flush with), and even if it was easy at our end, very few groups that could reproduce it have the spare capacity to ingest that. We do plan to publicly release it, but those plans have a lot of questions. Alternatively we could try to share summary statistics (as suggested by a post above), but then we need to figure out at what level is appropriate. In our case we have a relevant summary statistic of our data that comes to about 1 TB that is now far easier to share (1 TB really isn't a problem these days, though you're not embedding it in a notebook). But a large amount of data processing was applied to produce that, and if I give you that summary I'm implicitly telling you to trust me that what we did at that stage was exactly what we said we'd done and was done correctly. Is that reproducibility? You could also argue this the other way. What we've called "raw data" is just the first thing we're able to archive, but our acquisition system that generates it is a large pile of FPGAs and GPUs running 50k lines of custom C++. Without the input voltage streams you could never reproduce exactly what it did, so do you trust that? Then you're into the realm of is our test suite correct, and does it have good enough coverage? I think we have a pretty good handle on one aspect of this, is our analysis internally reproducible? i.e. with access to the raw data can I reproduce everything you see in the paper? That's a mixture of systems (e.g. configs and git repo hashes being automatically embedded into output files), and culture (e.g. making sure no one things it's a good idea to insert some derived data into our analysis pipeline that doesn't have that description embedded; data naming and versioning). But the external reproducibility question is still challenging, and I think it's better to think about it as being more of a spectrum with some optimal point balancing practicality and how much an external person could reasonably reproduce. Probably with some weighting for how likely is it that someone will actually want to attempt a reproduction from that level. This seems like the question that could do with useful debate in the field.
- 14 5y agoI remember in grade school reports seemed to logical. Propose something, create a hypothesis what you think you will see, record your data and what you observe during the experiment, summarize the results as to what actually happened vs what you initially expected. Most papers now seem like a foreign language and I can only glimpse at what is happening relying on some math genius to reply the significance.
- haihaibye 5y agoOne of the first suggestions I have is to use source control and store the Git hash of code used to generate data. A few times I've heard back "we don't have time for that" - pretty easy to see how the replication crisis flows from processes like that.
- jiggunjer 5y agoEven then, the entire software environment and even the compiler choice or different hardware could cause numerical differences.
- _Microft 5y agoPapers aren’t pop-sci articles, they do not target an audience that does not knows anything about the field yet. They are from experts for experts. If someone wants to familiarize themselves with the language, symbols and methods of a field, a textbook is a better thing to start with. Over time they will also learn the shared knowledge of the field that isn’t even mentioned in these articles.
- ketozhang 5y agoCertain scientific software packages (e.g., Tensorflow, pymc3, etc) do have frameworks that you follow to return pipeline and result objects that follow some data model that others can learn quickly (e.g., an arviz::InferenceData result object). I wish there was a more extensive framework where this is applied end-to-end from data input, to library components in a pipeline processes, to the result, and then to plot.
- fho 5y agoWorse in that a lot of researchers actually have only the slightest grasp on statistics. To the point that I would assume that a lot (1 in 20? :-)) papers will contain an error in their statistical analysis of their results.
- queuebert 5y agoAs a practicing scientist, I firmly believe the world would be much better off if we simply published version-controlled Jupyter notebooks on a free site, such as GitHub or ArXiv.
- throwaway984393 5y agoI think what you really want is a FOSS Mathematica. It's sort of code, but more convenient for just getting work done. You can pass around the whole thing (data + code). No need to learn software development skills or set up a development environment to replicate results; it runs everywhere. Plenty of power to do advanced things. Already a standard in research.
- robotresearcher 5y agoDo you do so? If not, why not?
- queuebert 5y agoI do share code that way, but the traditional ivory tower standards by which I am judged require "refereed journal publications" in high impact factor traditional journals. I'm trying to fight back against that, largely unsuccessfully. What would help me is to have the old geezers consider GitHub issues, PRs, and commits as a type of citation and to have a better way of tracking when my code gets used by others that is more detailed than forks. I also think citations of your work that find errors or correct things should count as a negative citation. Because otherwise you are incentivized to publish something early and wrong. Thus the references at the end of the paper should be split into two sections: stuff that was right and stuff that was wrong.
- d110af5ccf 5y ago> the references at the end of the paper should be split into two sections: stuff that was right and stuff that was wrong I've seen stuff like this said before but I don't think it would work. Most citations are mixed in my experience. A few objections, a bunch of stuff you aren't commenting on, and some things you're building on. Or you agree with the raw data but completely disagree with the interpretation. Others are topical - see <work> for more information about <background>. Probably more patterns I'm not thinking of.
- deleted 5y ago[deleted]
- dr_dshiv 5y agohttp://distill.pub/about/ http://distill.pub/about/ It’s been done and it is amazing. This is the best journal in the world imho
- antognini 5y agoIt is also on hiatus :( https://distill.pub/2021/distill-hiatus/ https://distill.pub/2021/distill-hiatus/
- otrahuevada 5y agoI myself enjoy reading papers. At least after I transfer them to a single column, 14/16pt tall font with real headers, that is. The graphic format itself is dated and annoying, yes, but I find the expositional tone and immediately searchable references pretty cool.
- shadowfox 5y agoHow do you go about converting the papers?
- dang 5y agoDiscussed at the time: Redesigning the Scientific Paper - https://news.ycombinator.com/item?id=16764321 https://news.ycombinator.com/item?id=16764321 - April 2018 (107 comments)
- csours 5y agoAnother way to say this is "The UX of Scientific Papers is Poor" It's not hard to imaging a better UX - for the field what are the top 5 questions you want to answer before you start reading? Eg: Sample size, Funding, etc. Put those at the top of the paper with symbols.
- d110af5ccf 5y agoSome medical journals have done something very similar to this for quite some time. They also define any acronyms used and have a brief summary of the key results.
- sam-2727 5y agoI think this is a great idea theoretically, but in reality for most papers I don't want to see the data/underlying code. While it would be great to publish data/code with the paper (in the field I've worked on the most, astronomy, most data is already published with the paper anyways), I don't want/need to look through a notebook with the underlying code of the paper in order to just read the intro/conclusions (and maybe one key methods section). Interactive figures are a great idea, but again, oftentimes I don't really care to interact with the figure, or fiddle stuff around, I just want to know why the paper is important and how I should use its conclusions. The two-column format of most papers is very useful for skimming. So instead I would argue notebooks shouldn't replace papers, but supplement them (as they sometimes do already, in fact, but perhaps journals could make it an actual requirement to create a supplementary notebook). As the article mentions, scientific fields are gigantic nowadays, and skimming papers is critical when you're citing 100+ references in your paper.
- deleted 5y ago[deleted]
- TuringTest 5y agoThe point of interactive notebooks is not seeing and having access to all the data - it's seeing the abstractions at work, having a direct grasp of how they act on particular examples as an aid to understand their formal definition. Nothing prevents you from having two-column notebooks, if you find that advantageous, as well as abstract and conclusions sections. The part that you don't get with static paper is that of navigating the abstraction ladder[1] up and down with direct manipulation aids, instead of having to work it all in your head or by following dense detailed paragraphs. [1] As also explained by Bret Victor in http://worrydream.com/LadderOfAbstraction/ http://worrydream.com/LadderOfAbstraction/
- GracefullyBlind 5y agoI think having the ability to focus on the things you care about the paper mostly is what would be more beneficial for all readers. You care more about an overview? You can easily find it (perhaps with graphics and walkthroughs), you care more about proofs? Then you can get them, what about code and experiments? And so on and so forth. Readability and scalability is about making all this data available in the publication record, but easy to navigate for whoever is looking for whatever.
- johnsutor 5y agoI feel like the website paperswithcode.com addresses this very well, especially with their feature "quick start in Colab". For example, here's the top paper on the website as of now: https://paperswithcode.com/paper/towards-real-world-blind-face-restoration https://paperswithcode.com/paper/towards-real-world-blind-fa.... Instead of going through the process of cloning a repo, initializing a fresh Anaconda environment from scratch, reading through nebulous, haphazard documentation about how to download the necessary training data, and then converting that training data into a format that's compatible with the code, I just click a link and run a couple of lines. Bam. I have an intuition about the code that's 100x better than reading the paper alone. Even though Colab isn't applicable to all fields and is largely used by the Data Science and Computer Science community, it is a promising step at modernizing science, especially the replicability of discoveries.
- musicale 5y agoIt's really disappointing that technical societies like the ACM and IEEE haven't done this already. For many journals and conferences there isn't even a way to submit the code or other digital artifacts with the PDF. A few have badging for whether digital artifacts are provided and whether the results have been reproduced or repeated by others - steps in the right direction at least. As much as I intensely dislike their practices of overcharging for journals and milking digital library subscriptions to fund administrative overhead, the technical societies are technically non-profits and exist to serve their members and the research and professional community. This is really something they should be doing.
- d110af5ccf 5y agoIt seems like you should be able to include more or less arbitrary binary artifacts as a form of supplemental information provided the peer reviewers skim and sign off on it and it fits within some size limit.
- pwr-electronics 5y agoI don't know if they necessarily should. It's easy enough to post your own supporting materials online. But it doesn't fit the publisher's mission of keeping a timeless paper trail of ideas.
- a-dub 5y agowhat came first? notebooks in mathematica or knuth's ideas on literate programming[1] ? regarding notebooks themselves, i feel like they're a high concept idea but i've yet to see them really click for me in practice. i find the small cells for code to be extremely unergonomic and that the interspersal of code and plots to be distracting from both the code and the plots (although pretty fantastic for demonstrating high level features of a library, programming language or environment). on a more fundamental level, i completely agree that mathematical notation is lossy, and that it takes a lot of skill to go from some arcane notation to an actual sense of what the relationships are- but, it requires no specific functioning technology to do so. i can review a paper from 100 years ago and understand it, where running a computer program from 20 years ago can be a challenge at best. i think that additional high touch experiences for data exploration and teaching are fantastic ideas, but i also think that maybe the base level of communication should be kept simple; both for the purposes of maintaining accessibility and history. where the linux kernel developers insist on 78 column listservs, maybe scientists should insist on camera ready documents when it comes time to share. i think that everyone agrees that better science would come from full data and code being supplied with publications, but interop is quite difficult as-is keeping code alive. i suppose the big question is: does it make sense to move science towards how software is done, where every bit of code is actively maintained over the years to avoid code rot, or does it make sense to come up with a scheme of freezing and archiving computing environments used in science so those in the future may be able to reproduce results or errors as they see fit. (something like, every paper must ship with a vm image for a widely available architecture that includes no proprietary code and all data used for results) interesting questions. how to fundamentally change scientific communication such that it is enriched with data and code properly is a harder/organizational problem that i think many have tried to solve (not to mention how this ties into another problem in science- idea validation/replication and knowledge rot). building software systems for exploration and data analysis (ie; computer as partner in exploration) sounds much more fun and likely to produce useful results! [1] https://en.wikipedia.org/wiki/Literate_programming https://en.wikipedia.org/wiki/Literate_programming
- multilogit39 5y agoAs an academically employed scientist of 20 years, the notion that scientific communication suddenly needs better standards puzzles me. The core research curriculum of nearly every scientific field I’ve seen, STEM or otherwise, is that the data needed for replication are non-negotiable. A paper that doesn’t include it would be table rejected by any editor. Or one would hope. This is taught at the UNDERgraduate level, for heaven’s sake. The thought clusters emerging from the recent “replication crisis” are a fascinating rabbit hole to crawl into. If you stay near the surface, you will find mostly young scholars cheerleading open science as the obvious solution to replication difficulties. The concepts of pre-registering your study, committing to sharing data, and publishing online are all various components of this idea, varying in their necessity by the author’s devotion to their cause. But there are several downsides to such a system that aren’t immediately obvious. For example, does the skill set of the successful scientist broaden to include how skilled they are at poaching ideas from public data that wasn’t immediately seen by their authors? Some of the more recent criticisms invoked the spectre of “platform capitalism”, and suggested the Facebook and Linkedin-ification of science by dumping all its data on a centralized platform would likely have a net negative effect. This article was written in 2018, and most of the discussions I’ve read since then have suggested that the open science initiative has failed despite the rapid penetration of Jupyter and visualization tools in the scientific process. Perhaps, like most things, the unseen market will pick and choose the good out of the dubious.
- BeetleB 5y ago> The core research curriculum of nearly every scientific field I’ve seen, STEM or otherwise, is that the data needed for replication are non-negotiable. A paper that doesn’t include it would be table rejected by any editor. Or one would hope. This is taught at the UNDERgraduate level, for heaven’s sake. This may vary based on discipline, but in both the subdisciplines of experimental and theoretical physics I was involved in: No - very few will provide the data/derivation. My professors were very open about this: They don't want to lose their competitive edge. Almost no experimentalist I knew could take papers from his/her field and reproduce the results, because the papers lacked enough detail to do so. They would mention a technique, but there are lots and lots of nuances involved when building equipment to carry out the technique[1], and these are intentionally excluded. It's unlikely you'll be able to build the equipment the same way the original authors would. [1] Most experimental physics involves building your own equipment, or at the least modifying existing equipment.
- GracefullyBlind 5y agoGreat read. Taking from my field (CS), I think a lot of papers suffer from the idea that you only are supposed to show "the interface", like, what the result is, what you achieved. The "how" or the "why" are sometimes neglected, regarded merely as a "technicality" to account for the rigorous mathematical framework that "must be there". There is little effort in making your results understandable and easy to replicate. Academia values paper production, which requires convincing peer reviewers that your results are not trivial and are worth publishing. Contrary to what the essay states, I don't think many scientists today think their research is "incremental". In fact, this word is used in many places as a derogatory term to indicate certain result doesn't contain enough novelty to deserve publication. Researchers are more incentivized to make their constructions and results as complicated and less accessible as possible. This is not just a theory, this is something I've seen over and over throughout the years.
- williamkuszmaul 5y agoIn my field, at least, I think the problem is less about the medium, and more about the incentives. Researchers are incentivized to write papers that seem impressive (and intimidating) rather than clear and intuitive. To make matters worse, this is an evolved trait: researchers whose papers are intimidating are more likely to succeed, which means they're more likely to have future PhD students, which means that the style of writing is more likely to get passed on. I think the main way to address this is to change the incentives. In particular, by creating publication venues that value simplicity and clarity (one such conference is SOSA, which has had a lot of impact on theoretical computer science in the last few years).
- diognesofsinope 5y ago> Researchers are incentivized to write papers that seem impressive (and intimidating) rather than clear and intuitive. Ah, a fellow economist lol. Lack of clarity is a strategic advantage because (1) (as you said) it looks impressive and (2) it's hard to validate that it's correct. So many papers contain such elementary statistics mistakes such as survivorship bias, e.g. 'returns to education' is almost exclusively measured by asking individuals who graduated (on average 50% of enrolled students don't) and respond back to surveys (good chance of bias). Pubs are how you get jobs. It's not about science anymore, it's about navigating bureaucracy for an elite job.
- deleted 5y ago[deleted]
- fho 5y agoThere is the whole WEIRD participants thing ... of which I am guilty too, only because that's the crowd you can easily recruit for experiments on campus.
- civilized 5y ago> Ah, a fellow economist lol. Lack of clarity is a strategic advantage because (1) (as you said) it looks impressive and (2) it's hard to validate that it's correct. Reminds me of the old "there are two ways of constructing a software design: One way is to make it so simple that there are obviously no deficiencies, and the other way is to make it so complicated that there are no obvious deficiencies."
- jimslime 5y agoWhy is this surprising? When all truth is subjective no truth matters.
- xipho 5y agoIt's easy to say things are obsolete when you are your own publisher (Victor, Wolfram, Perez) or when can suggest your favorite (even if it's very cool) approach (Jupyter Notebooks) as a potential key solution. If you can't be your own publisher, it's a much more difficult proposition. We're trying to figure out how to facilitate taxonomists publishing their own taxon pages, i.e. species descriptions, from a science 250+ years old. Our MVP use case is ~20k pages, one per species, for one project. There will be many of these projects, though maybe not many with 20k pages, and some with much greater than 20k pages. Updates are needed with as little latency as possible with data from from multiple sources. There has to be basically zero cost to serve these (I know, nothing is free). Sites must be trivially configurable (e.g. clone a GH template repo and edit a YAML file and some markdown). Even if we can get this in infrastructure in place we then have to figure out how to get the social structure in place to have this type of product recognized as equivalent to traditional on paper publishing, i.e. advance people's careers because they "published". In my field, until we give people the power to publish on their own, I don't see traditional publishing go away. Many in the past have indeed published (traditionally) their own species descriptions on their own dime, meeting the rules within the various international codes of nomenclature. I also don't have a problem with dead wood- if we go digital too fast we will loose so much for any number of reasons associated with the ephemeral nature of electron-based infrastructures.
- btheshoe 5y agoSimilar things can be said about the textbook and the lecture
- agucova 5y ago> Similar things can be said about the textbook and the lecture If you haven't yet, maybe look at Andy Matuschak's "Why books donʼt work" [1] and " How can we develop transformative tools for thought?" [2] which connect these same ideas to education in general. [1] https://andymatuschak.org/books/ https://andymatuschak.org/books/ [2] https://numinous.productions/ttft/ https://numinous.productions/ttft/
- janeroe 5y ago> Why books donʼt work They do. First, just like with everything else there are brilliant, good, mediocre, and outright poorly written books. Among those some may work for you, others miss the mark completely depending on your prior experience and background (as the author rightly notices, books are just a medium). Second, did the author expect to become a domain expert after finishing a single book? Clearly, his expectations are unrealistic then. You start somewhere, then use references to deepen your knowledge. That's a task requiring interest and dedication, but no "several lifetimes of research" (of mnemonics and learning methods) as he puts it, will replace that.
- agucova 5y agoI'm uncertain if you read the entire article, but the point is that even the greatest non-fiction books aren't optimal when the objective is learning. His expectations aren't unrealistic, the methods he's suggesting have decades of research suggesting to learning that's much more efficient than traditional reading.
- janeroe 5y ago> even the greatest non-fiction books aren't optimal when the objective is learning Says who? I say he's doing it wrong. > the methods he's suggesting have decades of research suggesting to learning that's much more efficient than traditional reading Where are the results of that research then, the culmination in the form of medium superior to books? His mneumonic quantum book doesn't look like one. People need understanding, not memorization.
- subroutine 5y agoWhile I agree with the premise of this article - that the scientific paper is obsolete - the proposed alternative does not address the primary inefficiencies with our current system. Dynamic 'papers' with illustrative examples are surely an improvement over current manuscripts. However, a flashy paper should be well down the list of goals for a research program. At the top of the list should be to make an important discovery or breakthrough. Currently there is way too much trivial shit being published in the ever-expanding number of scientific journals. Our current model for career advancement in academia is partially to blame for this. Landing a tenure track job requires having a prolific publication record. And once you've landed that coveted ladder-rank position, the pressure to publish only heats up, with tenure on the line. And then even after you've secured tenure, career advancement still depends heavily on the ol' publication record. Pre-tenure the pressure to publish in high-impact journals is immense; it's the kind of pressure that drives otherwise honest people to consider partaking in fraud (and unfortunately those who stick to their morals often lose out to those who fabricate results to some degree). Post-tenure the pressure changes from securing high-impact papers to just getting on whatever papers you can. Review boards for associate professors would more readily give a promotion to someone with 20 meaningless papers over the last three years than someone with 2 papers in CNS journals over the same timeframe, even though a single paper in Cell/Nature/Science is typically more impactful than 100 papers in Frontiers or other similarly dogshit journals. So post-tenure pay raises are based on getting as many words in print as possible, using the least amount of effort to do so. I almost forgot where I was going with this... right, so, in my opinion we need to switch from a publication-based mindset to a discovery-based mindset. We (the public) provide the NIH with $30 billion dollars per year, with the idea that such an investment will lead to medical breakthroughs, discoveries, and other innovations that can concretely improve our health and wellbeing. However, so much of that is wasted on the idiosyncrasies of career advancement inside the ivory tower. If I were the director of the NIH my sole purpose would be to end this nonsense. And my first order of action would be no longer accepting grant applications from individual PIs. I would only entertain grant applications from a small force of scientists (4-8 lab equivalents) with thorough and cogent plans for making breakthroughs on cancer or heart disease etc. I would change the minimum R01 funding amount from <$500k to >$10 million dollars. I would change the grant renewal timelines from every year to every 5 years but require yearly progress updates to ensure the proposed experiment were soundly conducted. I would encourage the equal reporting of both positive and null findings. Performance would not in any way be based on positive findings, only that the experiments were carefully run. I would require that all raw data be deposited into publicly accessible repos (not at the end of the study, not yearly, but whenever data is generated it should be made accessible asap). I would encourage that research groups bother with drafting/submitting interim manuscripts (interim meaning prior to completion of the full 5-10 year study) only if they find something important, otherwise just provide a comprehensive writeup at the end of the study. This final writeup would not be submitted to a 3rd party journal. It would be posted directly on the NIH website I'd have created for such results reporting. Naturally this would also be publicly accessible. That would be a start...
- imranq 5y agoThis is a little exaggerated. Most papers have to be somewhat readable to be accepted into journals and notable conferences. The fact that the layman cannot understand an advanced biology paper is nothing new. I'd wager the scientific paper "golden age" the author cites as having such readable papers were not very readable for the general public of the time. It's just we are taught those things in elementary school and so can see the concepts in those older papers much better than the people of the time can.
- whatshisface 5y ago>The earliest papers were in some ways more readable than papers are today. They were less specialized, more direct, shorter, and far less formal. Calculus had only just been invented. Entire data sets could fit in a table on a single page. What little “computation” contributed to the results was done by hand and could be verified in the same way. This is not even close to true. Look up Tycho Brahe's observations, and he's only the earliest I can think of.
- popcube 5y agoin another way, if you read the early article in biology, they are very easy to read and funny! e.g., Wallace's article looks like travel note. but use them in research just a nightmare...
- mitchbob 5y agoArchived: https://archive.ph/wXqPh https://archive.ph/wXqPh
- serverlessmom 5y agoI definitely agree with the author that a major gap that has continually failed to be bridged between the scientific community and the more general bulk of people is true understanding of the field of science as a whole. Even considering the way that most people casually throw around the term "research" is elucidating of this problem: a genuine lack of understanding of not just complex scientific and mathematical models but science as an industry and a tool for understanding life on this planet. None the less I wonder if the goal should be to make it so any person could understand a complex paper? Should all people strain to understand every study? There are experts in certain fields for that very reason. It is not always possible to accurately explain higher level concepts to people who lack foundational knowledge that can take years to accrue. I am not certain if changing papers to be more interactive is going to bridge the gap as this author hopes, or if it is even the goal that should be pursued.
- Ice_cream_suit 5y agoSo I am speed reading this: DNA Methylation and Protein Markers of Chronic Inflammation and Their Associations With Brain and Cognitive Aging https://n.neurology.org/content/97/23/e2340.abstract https://n.neurology.org/content/97/23/e2340.abstract I read the abstract, in reverse order. Discussion and Results first. If it seems plausible ( much research is trash, churned out to pad a CV ) and interesting, then I scan the Methods section. This heuristic helps classify 95% as utter trash or outside my current area of interest in less than 10 seconds. If it seems seriously interesting, I then read the full text the same way ie: backwards. At no point do I want pretty visualisations made by wannbe PhD candidates, full of misinterpretations, wishful thinking or outright fraud.
- ketozhang 5y ago> At no point do I want pretty visualisations made by wannbe PhD candidates See Figure (5). Your argument doesn't really counter any part of the scientific notebook. A notebook will still have the abstract and conclusion (result & discussion). The tools mentioned in the article describes how to restructure the methods, data, and figures. You're note going to look at these anyways until the abstract and conclusion intrigues you.
- tpoacher 5y agoI don't disagree with the main point of the article, but I think it underplays the extent to which most publications being information-poor is the fault of the medium rather than the low standard of writing that we've become accustomed to and complacent with over the years. This is not too unlike maintainable code. The code platform itself matters to some extent, but far less than the extent to which the author wrote with maintainability in mind.
- bborud 5y agoYears ago I did a lot of original research in a field that wasn't very well developed at the time. However, that research took place in the context of a startup, not academia. I was rewarded for producing solutions that worked - not for publishing papers. I was approached by a few academics about publishing what I'd worked on, but I never did. I never did because I did consume large stacks of papers every month, and I absolutely hated the pompous, obfuscated portioning out of ideas fragment by fragment. It was an unnecessarily time consuming, and often quite useless way of sharing information. Especially since source code often wasn't part of what was published so a lot of important information got lost (which I guess was the entire point of not publishing code). I particularly remember a 4-5 page paper that was so poorly written it took me a couple of readings (weeks apart) to realize that it described something I too had worked on. How bad is a paper when it is so obfuscated that it takes effort to recognize something you have worked on too? I wasn't interested in wasting time dressing up my notes in drag. And if my notes as they were were not good enough, well, then someone else would surely do the same work independently and publish something at some point. Lots of the things I worked on inevitably were described by other people. I have a love-hate relationship to scientific papers for the simple reason that they sometimes aren't really about science, but about scoring points in academia and certain types of research organizations. Yes, a lot of interesting goodies are published, but my god there is a lot of garbage that gets published. Not least because people in academia are incentivized to get as many papers as possible out of what ought to be a single publishable unit. If we incentivize authors to spam us, they will spam us.
- tpoacher 5y agoIf it's any consolation, even if the source code was provided, it is highly unlikely that it would have been of much better organisation or quality than the paper itself.
- periheli0n 5y agoAs a publishing scientist myself I would have hoped to read more about how I can actually publish Jupyter Notebooks in a way that is recognised academically. That's at least what the title implied for me. But the article is actually about Mathematica vs. Jupyter notebooks. Still, it's well researched and very interesting. Nevertheless, the question how to publish better remains open. I for one think that some progress could already be made if ArXiv published html articles by default, rather than those unwieldy PDFs that really only work best when printed on paper.
- jillesvangurp 5y agoThe paper as way to publish information is rooted in a world that no longer exists. That world consisted of scientists accessing information in printed form through libraries that subscribed to relevant magazines, journals, and what not from publishers. That's still a thing as far as publishers are concerned but I rarely set foot in a library after Google became a thing last century. The last thirty years have changed that game to basically no longer involve printing (other than for a very few select publications) and basically switching to a digital only publishing form. I completed my phd, read thousands of papers (well skimmed mostly, it's a fine art to zoom in on the relevant stuff), all without visiting the library more than once or twice. I only printed the ones actually worth reading in detail. And these days I consume vast amounts of information on screen without ever using a printer. My printer is fifteen years old and I just installed the second toner cartridge I ever bought for it. Yet we still pretend to have "journals" like we're in the 19th century. It's the equivalent to writing your friend a letter to inform them that your train is delayed by 5 minutes. Most sane people use some kind of instant messaging tool for that. Writing letters of course used to be a primary way to communicate for scientists. That too has stopped being a thing. People use email now. The whole point of publishing is to convey information in a form that's convenient to the reader and to solicit endorsement from your peers (via peer review). Peer review used to be implied by virtue of an editor choosing to select a certain paper for publishing. That in turn implies they would have consulted a number of peers about the suitability of that paper. It's sort of the super tedious equivalent of a soliciting a thumbs up button in a social network. If you publish on linkedin because you are some kind of wannabe influencer you basically need to get people to 1) read your stuff and 2) click the like or share button. Scientific publishing basically is not that different. You have wannabe scientist that want to get the attention of the influencers (reputable peers) so people will be convinced they know their shit. This ultimately translates into degrees, research funding, and tenure track positions. The whole process is kind of biased towards metrics because that's how universities choose to allocate their money. An ambitious scientist behaves basically in a similar way as a linkedin influencer and will try to game the system by flooding the system with a lot of content and getting their buddies to sign off on it. There are a lot of mediocre articles that get published in obscure places with cliques of scientists basically doing each other favors by referencing each other's work; or worse self referencing. In linkedin terms, this would be the absolute drivel that nobody likes that gets re-shared by a few people that also don't manage to produce much content of interest. So, here's a thought, maybe get this a bit more out in the open and give scientists some modern tools to endorse each other's work. The best endorsement is a reference. A link basically. Tracking links between bits of paper is super tedious. These papers need permanent URLs. And they need to be digitally signed by their authors so we can have some authenticity and prevent cheating. And scientists need a place where scientists can debate and exchange thoughts about these papers. That used to be a big tradition between scientist back when they still wrote letters to each other or used journals to criticize each other's work. Curating and aggregating work by means of linking to it is a job that should not be reserved for fussy editors of non paper based journals that absolutely nobody ever reads cover to cover. HN for science; why not? Why not have a multitude of websites referring, editorializing and commenting on published work? How is that not a thing?
- agnosticmantis 5y agoComputational notebooks are great, but only seeing what the author (or coder) wanted you to see is not enough to evaluate their work. In addition to data and code, we need to see the path they took, the exploration and experimentation that led to the final presentation of ideas in the notebook. For this we can use cloud based environments controlled by funding agencies/universities that ensure every interaction with data is recorded from the very beginning. Something like this would at least reduce the risk of p-hacking practices that would otherwise be there even if everyone used notebooks instead of papers.
- deepzn 5y agoJames Somers is one of my favorite science/tech writers. Thank you, James for enriching my mind over the years. I have a bunch of your New Yorker articles to read but throughly enjoyed your 'The Coming Software Apocalypse' story. Guys check it out here https://www.theatlantic.com/technology/archive/2017/09/saving-the-world-from-code/540393/ https://www.theatlantic.com/technology/archive/2017/09/savin... And he has all his articles listed on his site- https://jsomers.net/ https://jsomers.net/
- beckman466 5y agowow great stuff, thanks for posting!