4 ms·
Science Is Open Software
- jegp 13d agoTL;DR I claim that modern science is synonymous with open source software. This post explains why, why it matters, and what you can (and should) do next.
- gradus_ad 13d agoWhat's the equivalent of closed source?
- jegp 13d agoEquivalent? In the analogy of the math or physics results it would be a mental model in someone's brain that you can't access or verify. You just hope it's true
- IanCal 13d agoOpen source is much more than source available though. Its about licensing.
- altmanaltman 13d agoBut that doesn't map to software at all right?
- jegp 12d agoI think the idea does: the point is that closed-source software isn't useful for anyone else than the (copyright)owner. In scientific standards, that is. Similar to esoteric and cryptic theories that resist accessible explanations.
- random3 13d agoIt’s research happening privately without publishing, usually going into products
- Matumio 13d agoLike if CERN published the discovery of the Higgs boson with 99.9997% certainty, but refusing to tell you how they calculated that number, or what equipment they used and how they calibrated it, in order to prevent other labs from copying their methods. Or like a machine learning lab claiming SOTA on a benchmark, beating a well-known method that they re-implemented, possibly with bugs, on their private dataset, for millions of compute. But you don't get the source to check, and they don't release any intermediate results or ablation experiments. Aka, from the outside you can't distinguish it from corporate marketing.
- jibal 13d agoYou argue that open software is science, which is not at all the same as claiming that science is software. ("is" in this context is not equivalence -- "a poodle is a dog" != "a dog is a poodle".)
- jegp 13d agoI agree the post is muddy about whether the relationship is bijective (equivalent, poodle=dog) or injective (onto, poodle is a dog). I make it slightly more precise in the statement "I posit that open source software is a necessary condition if we are to science in a computerized world". That's where the "is" comes from in the title. Throughout history, this definitely has not been the case. I'm arguing that's changing.
- jibal 13d agoIt's not just "muddy", it's thoroughly inconsistent and impenetrable (which probably has a lot to do with why there is so little engagement here). You say "TL;DR I claim that modern science is synonymous with open source software" which is radically different from your "slightly more precise" statement. I won't put any more time into this ... good luck in figuring out what it is you really want to claim and presenting a coherent and cogent argument for it.
- jegp 13d agoThanks for the well wishes
- random3 13d agoScience is open, but science is not software and software definitely not science.
- jegp 13d agoDid you read the post...?
- jibal 13d agohttps://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html > Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".
- throwaway27448 13d ago[flagged]
- jegp 13d agoArguments are fine. Vacuous statements need grounding. The follow-up is way more detailed
- random3 13d agoYes. It conflates a bunch of things > TL;DR I claim that modern science is synonymous with open source software That's a strong statement that's not supported by the arguments and IMO misguided. I don't have a problem with "open", but rather with "software". Both science and software deal with models, however the focus is quite different. I suspect you conflate theory with models. The goal of science is to produce and test theories — that's an inductive/abductive process. A model, regardless of whether it's reified into mathematical formulas or software, is a means of making a theory operational enough that its consequences can be derived and confronted with observations. Software often starts downstream of this: it's a reification of theories, models, algorithms, or findings that are the result of research. Of course software can also be used as part of the research process itself. The distinction is roughly the familiar one between research and development.
- willtemperley 13d ago> modern science is synonymous with open source software. Another problem with reproducibility is the openness of the underlying data. Many academics are terrified of giving away the golden goose and the software is often useless without the data. However many scientists do work openly, e.g. The Journal of Open Source Software: https://joss.theoj.org/ https://joss.theoj.org/
- setopt 13d agoJOSS is great, I’ve both published with and reviewed for them, and enjoyed it more than traditional journals. The process felt more constructive than destructive, in a sense. I believe they’re always looking for new volunteers to review papers, so please do volunteer if you are able.
- jegp 12d agoI've seen quite a few academics wrangling with patents. What should come first? Proliferation of the sciences or (potential) profit? Your point about the golden goose is pretty interesting and others in this thread has pointed to the misaligned incentive structure in academia. I'm curious, do you think things like JOSS could help open up data and "the golden goose"? I'm not sure the funding bodies I'm interacting with would respect this kind of initiative, but maybe it's just a matter of time.
- willtemperley 12d agoMy experience in academia was that data management is not well respected. Universities often see data as part of the library system which leads to a disconnect between data and research. I’d suggest data management needs to be a well paid career in academia, with senior level influence to enable long term open access to data. I was on a grant from the Wellcome Trust putting an epidemiology database online. This remained available until the PI moved on and their successor was decidedly less public spirited, so sadly it’s not available anymore. What can I do about that when I’m seen as a technician? I have over 8000 citations from my database work but there’s really no career in academia for me.
- throwaway27448 13d agoWe really need to ban tech people from using the word "open".
- jegp 13d agoErm. Why?
- throwaway27448 11d agoIt stopped having meaning a few decades ago when "open source" started to be used interchangeably with "free, libre, open source software". It's all vibes now—did you release anything aside from a strictly commercial offering? well, it's open then!
- samayashar 13d agoNice read. I believe that traditional software is a great way to showcase the proofs when it comes to physics and mathematics. You can easily code up a theorem in a language of your choice and justify that 'Okay, the output matches the expected value'. I am particularly fascinated by labs like DeepMind [https://deepmind.google/science/ https://deepmind.google/science/]. The recent advances in their frontier models that are able to predict diseases before they're diagnosed is incredible. This is what AI should be built for and actually do!
- jegp 13d agoThanks! I appreciate that. Your point about DeepMind and frontier models is spot on. When they "embody"/build on the science done before them we get absolutely mindblowing synergies. But I wonder what happens when the LLMs become way smarter that us: why even loop us in? I guess that's related to the recent field medalist letter https://mathandai.org/ https://mathandai.org/
- flopsamjetsam 13d ago> Every result is instantly reproducible. When you read a paper claiming that a new drug reduces symptoms by 30%, you click a link and watch the exact analysis run in your browser. The data processing, statistical tests, and visualizations execute in seconds using the same environment the authors used—preserved perfectly through reproducible containers. At least some journals have this as a stipulation e.g. https://www.nature.com/nature-portfolio/editorial-policies/reporting-standards https://www.nature.com/nature-portfolio/editorial-policies/r... Particularly the "data availability" and "Availability and peer review of computer code and algorithm". However, in my limited experience, of trying to reproduce certain scRNA-seq processing pipelines, in practice it's never available as just a Github link. I can understand that some/many researcher's code is not in good shape, so I think it'll be quite a stretch to have this available. I do think it's laudable though, to try and make it available. It would certainly have been very useful for me in the past.
- cge 12d agoI try to do something like this with my publications, and encourage others to. My goal is to have the pipeline from raw data to complete figures and manuscript in a repository, with cached data for computationally expensive analysis and for stochastic simulation results, and the option for the user to just use those or run the full pipeline, with or without the same random seeds. I just make clear that the code was run-once code and is going to be messy compared to code refined over time and diverse uses. I generally use Zenodo to a GitHub repo, however, in case GitHub decides to do something bad in the future. Making sure things run far in the future can also be a challenge. Sure, you can use a container: will the base of that container be available in 30 years? And with that said, for experimental work, this approach does not make things fully reproducible; it only makes the analysis reproducible. There are always factors that influence experiments: research is by definition at the edge of our understanding, and reality has countless variables, including ones no one has thought of, known about or thought important.
- amarcheschi 12d agoYou're goat, I'm trying to reproduce code from a paper and by following their instructions I can't even get packages to install because they conflict
- marsven_422 13d ago[dead]
- D-Machine 13d agoScience should be more like this, in current times, yes. But until much of academia is burned to the ground, or until science can be properly separated from modern academia, this will never be so. The current academic incentives are all wrong: low-quality research is rewarded and results in publications, whereas high-quality research (that takes time, and usually reveals that most exciting publications depend on p-hacking or other highly data-dependent analyses and selective presentations) is not published or actively blocked during peer review. So instead you get BS arguments about how data can't be released for various privacy concerns (when in reality the vast majority of most datasets are trivial to scrub of identifying factors, and even in more complex datasets where you need to consider k-anonymity, it is still trivial to release data that allows replication of core analyses), and academic science is increasingly irrelevant unless it is tied to tech and industry, where producing junk actually has real negative economic and personal consequences. I don't know what world this article / post lives in, but it isn't the messy world of actual reality.
- stalfie 13d agoHear hear! There are so many obvious improvements to how almost everything is done. For instance, in medicine review articles as a class of articles largely represent a giant waste of time. RCTs flatten all their gathered data during publishing, summarizing complex trial data, which is gathered but never published, into a few numbers. Then review articles take a bunch of flattened data, discard the articles that don't fit the exact question they are reviewing, and then publish a doubly flattened conclusion. If any of the included articles turn out to have flaws, if treatments change in retrospect, if you are looking for the answer to a slightly different question or you are looking at a different subgroup, then the review is useless and has to be repeated. All of these tens of thousands of man-hours could be replaced by a few GitHub repos, if only RCTs would just publish their damn data. Then you could just run and rerun the statistics on whatever subgroup you're looking for, instead of combing through decades of review articles answering slightly different questions, looking for the answer between the lines. With LLMs making mining of large scale datasets almost trivial (with the process most likely becoming trustworthy within a few years), the current status quo is looking more and more antiquated. If you want to be even more radical, hospitals could just publish their data continuously. Of course, it is easy to point to the risks of doing so, but what's often ignored is the benefits. It is hard to overstate just how many medical mysteries a hospital encounters on a daily basis, how much unknown we are navigating in practice. The current norm is that 99.99% of these cases are never published, and are only ever thought about by a small group of people who happened to be at work. Particularly, when someone dies of something no one figured out, it is never published anywhere, because even if you tried it is not interesting reading material for a journal to publish. And no one ever tries because they're scared of being called out for a mistake. A hospital is essentially a continuously running and extremely interesting experiment, where 99.99999% of all results are thrown in the garbage, and the only published data is subject to extreme selection bias. All of this could be different, and the risks involved are actually quite small in practice. It is easy to automatically anonymize data quite well, but extremely difficult to absolutely guarantee that it is anonymous. And since current ethical norms are extremely averse to any degree of risk, and usually entirely ignore potential benefits, we all suffer for it. It is not entirely unlikely that someone reading this post will one day die because of something that could have been prevented, had things been different.
- sarfaraznaushad 13d ago[flagged]
- txrx0000 12d agoI agree with the general sentiment, but there's one major caveat. We should implement reproducible programs on top of a virtual machine spec like JVM or WebAssembly rather than replicate the entire environment. It's more practical to do and doesn't push software towards further centralization. Let people use whatever OS and VM implementation they want, or even write their own.
- jegp 12d agoIn principle I would agree. But, on a more philosophical level, couldn't you make that same argument about C? Or even assembly? Or even digital computers? Less facetious, it seems to me that the particular abstraction is less important. As long as it's unambiguous and widespread.
- txrx0000 12d agoI think there is a preferable category of abstractions that are adequately unambiguous, hardware-agnostic, and only occupy a thin layer of the stack. The argument could sort of be made for C, but not native assembly or digital computers because that would push hardware towards centralization. A useful way to think about this is to consider existing human languages like English or mathematical notation. When you read these English words, it's like you're executing a program I've written that will change your brain state. A pattern in my brain is encoded, stored, and transmitted in a language that we've both learned in our own separate ways, then decoded to a pattern in your brain. But your brain is quite different from mine, and you're free to implement any sort of sandboxing you want against this information, and even against the language specification itself, in your own head. Erasing this barrier is equivalent to erasing individuality, so I think we ought to be able to do this for digital programs as well. I hope that in the not too distant future, we will be able to design and manufacture custom hardware on a per-person basis. Because that's what it will take to preserve human autonomy and bodily integrity in the upcoming era of AI and brain augmentations. I don't want my neural interface to have a hardware-level backdoor (https://en.wikipedia.org/wiki/Intel_Management_Engine https://en.wikipedia.org/wiki/Intel_Management_Engine) like my desktop computer. There is also a bigger evolutionary problem behind this, and I expand more on that in this tangentially related past comment: https://news.ycombinator.com/item?id=49690354 https://news.ycombinator.com/item?id=49690354
- Muhammad523 12d agoReplace "Open" with free as in "freedom" gnu.org
- enbugger 12d ago> NixOS is quickly becomming the biggest and best tool there is. It will guarantee that your code will run exactly the same way, even 100 years in the future. Docker, Conda, and similar tools are better, but NixOS gives more comprehensive guarantees. I like how this is dropped as a fact. Dare to explain why though? Especially vs Docker. NixOS is not even standardized. No guarantees it will not be superseded by some descedant or eg. Guix in a near decade.
- btrettel 12d agoIn my work (scientific or otherwise), I try to avoid dependencies if possible. That's not always possible, so a solution like Docker or NixOS is needed, but the problem can be improved a lot without a technical solution. Either feels like fighting an uphill battle though as most researchers think short term and just pick whatever is convenient in the moment.
- jegp 12d agoRegarding the biggest: nixpkgs sits at around 140k packages, way more than others. Best: I still argue that Docker and Conda are more accessible, but my point was to go for reproducible, declarative science. Nix environments are exactly that. They're not perfect and are, as you point out, not standardized. But they cover much more ground that Docker. If you trust the upstream nix repo, you can get bit-level equality at every single build you (or anyone else) does. Docker relies on huge binary blobs that you can't inspect and that can be pretty much arbitrarily swapped around. I'm totally fine if Guix takes over. Or the next big thing. As long as it's declarative and reproducible.
- flimflamm 12d agoOne can decide to dedicate their own time to creating open SW. Typically someone (like tax payers or private companies) pay for the creation of scientific discoveries. Thus there is indeed a difference in the monetary intensives.
- shevy-java 12d ago> modern science is synonymous with open source software But why does the public have to pay for e. g. Elsevier? We pay for research of scientists already via taxpayers money (at the least in a civilized country), then we have to pay again for a private entity. If science is really open then it also needs to require public publishing. Gangsters such as Elsevier and others should not be able to drain the public here. Taxpayers financing something should also require public access to findings, at all times. Instead, Elsevier, Springer etc... get more public money while keeping things private. That's the antithesis to science.
- jegp 12d ago100%. I'm not saying the vision we're talking about is there yet, I just think it's an enticing thought. The existence of (predatory) publishers is a sad and miserable joke in its own right.
- btrettel 12d agoThe title reminds me of this, which is arguing the opposite direction: https://softpanorama.org/Articles/oss_as_academic_research.shtml https://softpanorama.org/Articles/oss_as_academic_research.s...
- runningmike 12d ago100% disagree! Do not mix and cherry pick terms and definitions to make your point. That's not scientific -) Take e.g. a look in https://opensciencemooc.eu/ https://opensciencemooc.eu/ and check a nice accepted handbook like "The Turing Way" [1] [1] https://book.the-turing-way.org/ https://book.the-turing-way.org/
- jegp 12d agoI'd love to hear you expand on how I cherry pick terms and definitions. The Turing Way seems great and I'm all for education. One of the most interesting conversation topics in this thread is, to me, how to create the necessary incentive structures. Know how is only part of the way. We need to secure the credit assignment for "openness" both in academia and industry.
- analog31 12d agoWhen all you have is a hammer, everything looks like a nail. I refer to what I think the author is asking for, as "push button reproducibility," i.e., the idea that the results will reproduce themselves at the push of a button, anywhere, at any time in the future. I have a couple of misgivings about this. First, the whole idea of "open" research predates computer technology. Forcing science to keep up with the latest ideas in software distribution is too much of a burden, when science is already too risky and slow. I had the odd privilege of learning the scientific method from my mom, before there was widespread access to computers. Her version was that a study should be reproducible by a reasonably skilled person. This is a greatly relaxed standard, but is realistic for a discipline that spans decades if not centuries. I supplied all of the data and code for my thesis research (and a sufficient number of mechanical and electrical drawings). But nobody has Turbo Pascal today, and some of the commercial instruments such as specialized lasers were already obsolete by the time I finished. Also, the experiment was dangerous, and might not pass safety review today. It required about $500k of equipment and a dedicated lab. Today I have the luxury of saying that if my code fails upon loading a new version of a dependency, the person who discovers that failure is probably skilled enough to fix it, and my work rarely hinges on the idiosyncracies of dependency versions. If it goes into a product, they'll totally rewrite it anyway. Second, science is still at its core an experimental discipline. Even in physics, there are more experimentalists than theoreticians. Reproducibility means roll up your shirt sleeves and head for the lab. To this day, some processes have not been mechanized, and you still need to spend years developing "lab hands" which not all people succeed at. I think there's a clue in the fact that the social and medical sciences seem to be the most deeply embroiled in the reproducibility crisis. It's because the quality of results depends on the the quality of measurements, and it's just harder when dealing with living subjects or one-of-a-kind specimens (such as the earth's climate). In fact, not much more than a century ago, it was believed that studying those things was beyond the reach of scientific methodology. Third, we're not going to stop doing science in areas where it's hard, particularly in medicine, but we're also not going to staff up in areas such as software development, to make science work better. People are suffering from disease right now so there's always an urgency to finding cures, plus an obvious profit motive. Disclosure: Experimental physicist, developing better measurement equipment.
- 12d ago
- 5555watch 12d agoGPTZero says: "We are highly confident this text was AI generated (100%)" Am I the only one that's tired of "my random shower thought turned into full article with AI" articles?
- jegp 12d ago:-) I can comfort you with the fact that the text is not AI generated. I did use AI to proof read it and I liked some of its suggestions to improve the flow of the text. English isn't my first language, so this is a great help for me.
- crustyoldhuman 12d agoThis article is correct but it bummed me out. It made me think of how much of science has been perverted into other goals, like medical science for example. The advancement of human health largely depends on corporate interests, and that is so insane to my brain it hurts to think about. very few independent scientists can research anything because everything costs money so you need a corporate interest to even do research. And the corporation gets the "rights" to those findings? It's absolutely insane to discover something natural and claim it as your own, and science is natural nobody is inventing it or being creative and writing it themselves, they're essentially walking up to a mountain and saying "Ok it's mine now I'll charge you 100$ to walk on the mountain". It genuinely makes my brain do somersaults in my skull that we've somehow backed SCIENCE of all things into this weird gatekept scenario its in now. It wouldn't bum me out so much but it clearly does nothing but hinder progress The point at the end "The scientific revolution succeeded because it insisted on transparency, reproducibility, and constant scrutiny." is a good one. But the real kicker that makes the situation so bleak is it's not the scientists who get to decide whether or not these things get applied to science or not, it's government policies and corporate interests.
- SR2Z 12d ago> It wouldn't bum me out so much but it clearly does nothing but hinder progress I don't know if this is true. Lots of this stuff is inherently expensive - clinical trials, research into vast numbers of compounds, and then mass-producing the drug are all things that cannot be feasibly done without at least millions of dollars. The US and China are the current world leaders in biotech, and it's because both of them have massive infrastructure to funnel billions of dollars into research. Yeah, the IP law could be reformed (I am a big believer in reducing IP protections in general) but the truth is that SOME FORM of protection is necessary to convince people with money to fund this kind of research. The government is simply not capable of this level of spending for such uncertain rewards; it doesn't have the proper incentives to recognize and promote good research while defunding useless research.
- pedalpete 12d agoAn interesting thought experiment, but I'm not sure I follow the reasoning of these two statements, which is the crux of why open software fits this mold. 1. Reproducible, meaning executable, as well as modifiable, and 2. Reliable, meaning that the results are consistently trustworthy Why does modifiable fit reproducible. Modifiable fits the open source theme, but not the science theme. Can someone set me straight on this if I am missing something? Number 2 is something I've been thinking about more recently, and the definition of science from the post being "testable hypothesis and predictable". Medicine is testable and predictable, but not reliable. To me reliability is the job of engineering. To take what is known in science and make it reliable. I used to work at CSIRO (Australia's science and technology org) and I remember my boss telling me that researchers always think their work is ready to go from the lab into the real world, but it is our job as engineers to make that translation so that the work outside of the lab. That's what I think about when I think of reliable. Thoughts?
- jruohonen 11d agoIn an utopia, yes, but I have raised the argument also elsewhere: who wants to be a maintainer of a paper forever? So, yes, science is indeed a lot like open source software, but more like the write-only software platforms and repositories are full of. Replication is also a lot more important than reproducibility in most fields and for most questions.