5 ms·
For some reason this struck me as inappropriate for the outlet. It's a nice piece as an introduction to array programming with numpy, but seemed out of place to
by teorema 6y ago
For some reason this struck me as inappropriate for the outlet. It's a nice piece as an introduction to array programming with numpy, but seemed out of place to me.
- ddavis 6y agoIf, going forward, 5% of all papers that use NumPy to get their results actually cite this paper, it will be one of Nature's most cited papers every year.
- lumost 6y agoThere's an interesting trend of what content gets published in peer-reviewed journals vs. blogs/github/etc. I suspect there is an audience segment that strongly values peer reviewed pieces that are equivalent content wise to introductory material in a variety of formats. I wonder if github should add a "Review" feature to provide a similar content authoring experience.
- clickok 6y agoIt would be nice if citing repositories were easier-- either for generating a reference for my own code or acknowledging when I've used someone else's code in my research. There's tons of math and physics blogs that contain useful results that the author wanted to make available but didn't manage to incorporate into a paper. I wonder if there'd be any interest in a sort of GitHub for proofs? It could even use git, since (assuming consistency) isn't math just a DAG anyways (and therefore isomorphic to a neural net, as are all things).
- throwawaygh 6y agoTraditionally that sort of stuff goes in tech reports, dissertations, or text books. What's missing is the dissemination piece. Somehow people will absolutely refuse to take seriously the job of citing code they use, even when their main result is obtainable by "and then I ran something from scipy/numpy/pytorch/etc."
- westurner 6y agoYou can get a free DOI for and archive a tag of a Git repo with FigShare or Zenodo. If you have repo2docker REES dependency scripts (requirements.txt, environment.yml, postInstall,) in your repo, a BinderHub like https://mybinder.org https://mybinder.org can build and cache a container image and launch a (free) instance in a k8s cloud. Journals haven't yet integrated with BinderHub. Putting the suggested citation and DOI URI/URL in your README and cataloging citations in an e.g. wiki page may increase the crucial frequency of citation. A Linked Data format for presenting well-formed arguments with #StructuredPremises would help to realize the potential of the web as a graph of resources which may satisfy formal inclusion criteria for #LinkedMetaAnalyses.
- cycomanic 6y agoThe issue is that none of the citation count engines (Google scholar, scopus, Web of Science...) count citations on those DOIs. So for a researcher who needs to somehow demonstrate impact through citation counts, it does not really help unfortunately.
- westurner 6y agoWe could reason about sites that index https://schema.org/ScholarlyArticle https://schema.org/ScholarlyArticle according to our own and others' observations. Google Scholar, Semantic Scholar, and Meta all index Scholarly Articles: they copy the bibliographic metadata and the abstract for archival and schoarly purposes. AFAIU, e.g. Zotero and Mendeley do not crawl and index articles or attempt to parse bibliographic citations from the astounding plethora of citation styles [citationstyles, citationstyles_stylerepo] into a citation graph suitable for representative metrics [zenodo_newmetrics]. bitcoin.org/bitcoin.pdf does not have a DOI, does not have an ORCID [orcid], and is not published in any journal but is indexed by e.g. Google Scholar; though there are apparently multiple records referring to a ScholarlyArticle with the same name and author. Something like "Hell's Angels" (1930)? No DOI, no ORCID, no parseable PDF structure: not indexed. AFAIU, Google Scholar does not yet index ScholarlyArticle (or SoftwareApplication < CreativeWork) bibliographic metadata. GScholar indexes an older set of bibliographic metadata from HTML <meta> tags and also attempts to parse PDFs. [gscholar_inclusion] Google Scholar is also not (yet?) integrated with Google Dataset Search (which indexes https://schema.org/Dataset https://schema.org/Dataset metadata). FigShare DOIs and Zenodo DOIs are DataCite DOIs [figshare_howtocite, zenodo_principles]; which apparently aren't (yet?) all indexed by Google Scholar [rescience_gscholar]. IIUC, all papers uploaded to https://arxiv.org https://arxiv.org are indexed by Google Scholar. In order for arxiv-vanity.org [arxiv_vanity] to render a mobile-ready, font-resizeable HTML5 version of a paper uploaded to ArXiV, the PostScript source must be uploaded. Arxiv hosts certain categories of ScholarlyArticles. JOSS (Journal of Open Source Software) has managed to get articles indexed by Google Scholar [rescience_gscholar]. They publish their costs [joss_costs]: $275 Crossref membership, DOIs: $1/paper: > Assuming a publication rate of 200 papers per year this works out at ~$4.75 per paper [citationstyles]: https://citationstyles.org https://citationstyles.org [citationstyles_stylerepo]: https://github.com/citation-style-language/styles https://github.com/citation-style-language/styles [gscholar_inclusion]: https://scholar.google.com/intl/en/scholar/inclusion.html#indexing https://scholar.google.com/intl/en/scholar/inclusion.html#in... [figshare_howtocite]: https://knowledge.figshare.com/articles/item/how-to-share-cite-or-embed-your-data https://knowledge.figshare.com/articles/item/how-to-share-ci... [zenodo_principles]: https://about.zenodo.org/principles/ https://about.zenodo.org/principles/ [zenodo_newmetrics]: https://www.frontiersin.org/articles/10.3389/frma.2017.00013/full https://www.frontiersin.org/articles/10.3389/frma.2017.00013... [rescience_gscholar]: https://github.com/ReScience/ReScience/issues/38 https://github.com/ReScience/ReScience/issues/38 [arxiv_vanity]: https://www.arxiv-vanity.com/ https://www.arxiv-vanity.com/ [joss_costs]: https://joss.theoj.org/about#costs https://joss.theoj.org/about#costs [orcid]: https://en.wikipedia.org/wiki/ORCID https://en.wikipedia.org/wiki/ORCID
- MayeulC 6y agoOwning to the distributed nature of git, and the properties of the hashes it uses, it is probably enough to put a full commit id in a paper to securely reference a software project, regardless of its hosting platform. We'd just need a dedicated search engine, and a way to automatically extract those from papers, to clone and archive repos.
- spappal 6y ago> the properties of the hashes [g]it uses Git uses SHA-1, a hardened version since 2017, and are now doing per-repo upgrades to SHA-256 [0]. Lots of repos are presumably still on SHA-1 (and users on older versions of git). As of 2020, chosen-prefix attacks against SHA-1 are now practical. [verbatim from 1] But I don't think second preimage attacks are practical yet. Linus Torvalds argued in 2006 basically that it's irrelevant whether git's hash function is second preimage resistant. Selective quoting: > remember that the git model is that you should primarily trust only your _own_ repository [2] > [a malicious] collision is entirely a non-issue: you'll get a "bad" repository that is different from what the attacker intended, but since you'll never actually use his colliding object, it's _literally_ no different from the attacker just not having found a collision at all [2] All that is just to say: git originally chose its hashes for the above mentioned "git model", thus didn't 100 % care about second preimage resistance. For your suggested search engine, depending on how the database is collected you might not be able to trust "your own repository" (if it's crowdsourced I could register another codebase with the same hash as Linux). A second preimage resistant hash function would be a requirement for the suggested use case. [0]: https://git-scm.com/docs/hash-function-transition/ https://git-scm.com/docs/hash-function-transition/ [1]: https://en.wikipedia.org/wiki/SHA-1#cite_ref-8 https://en.wikipedia.org/wiki/SHA-1#cite_ref-8 [2]: https://marc.info/?l=git&m=115678778717621&w=2 https://marc.info/?l=git&m=115678778717621&w=2
- Fishysoup 6y agoMaybe they're trying to promote a move from Matlab (which is close to exclusively used in a lot of academic fields) to Python, and not to Julia. Or maybe that's too far? Either way yeah I also find it slightly weird that this was somewhere in Nature.
- etimberg 6y agoYeah, this is a review article. Nature is not the right spot for it.
- flor1s 6y agoConsidering Nature labelled it as a review article and still published it, it does seem like Nature is an applicable venue for publishing a paper like this.
- joan_kode 6y agoNature has numerous review articles every year since 1974 (right on their website): https://www.nature.com/nature/articles?type=review-article&year=1975 https://www.nature.com/nature/articles?type=review-article&y...
- modeless 6y agoSeems to me like it's just a way for the journal and the authors to collect a gigantic number of citations to win the academic citation game. Not that there's anything necessarily wrong with that; NumPy deserves it of course.
- lacker 6y agoI think this is great. Nature is really a way for scientists to score points, not a publication that you read cover to cover that needs stylistic consistency. Right now the academic citation-count scoring mechanism doesn’t give enough incentive for people to work on the important infrastructure pieces like Numpy. So this is a good step towards putting scientific priorities in the right place.
- Aperocky 6y agoYep, especially when the author are the people who wrote numpy, I have absolutely no problem with that. It's about time they be recognized for their contribution to the tools of science.
- teorema 6y agoI definitely think things like the infrastructure don't get enough credit. I also mean no criticism of numpy. But is Numpy per se conceptually that innovative, from a computer science perspective? I guess to me this just seemed unusually introductory, about a specific library for a specific language. Put another way: if I was going to cite numpy, would I cite this? Probably not. Would I cite this paper for any of the more general concepts it covers? Probably not. I'd probably even argue someone shouldn't cite it for that latter reason, as those concepts supercede numpy (and appear in other languages under other names).
- konjin 6y agoI forget who, probably Hamming, said his most cited paper ever was just an intro to statistics for biologists. This is not a new thing and it's not a bad thing. Papers are meant to spread information. If a field isn't aware that another field has solved a problem of theirs than even a 101 level paper is worth writing.
- JimTheMan 6y agoWhile NumPy may be old news to the programming community, I have experienced a surge of programming capability into wet bio labs that was not there a decade ago. This kind of article heralds its adoption to mainstream biology. It's now well known enough to interest biologists in general!
- wenc 6y agoThat's a good point. At first I was like, NumPy's fame -- owing to the rise of Python in scientific computing and data science circles -- has gone far beyond that of most things ever published in academic journals, so this hardly seems necessary. But I see your point: Nature has always been held in high regard among natural scientists, and though most recently minted natural scientists have at least a passing familiarity with Python, the generations of scientists before them probably don't. Nature at least has enough of a cachet to grab their attention. Plus having a publication in Nature does open doors. Travis Oliphant and some of the already-famous co-authors probably don't need these doors opened, but I'm sure others on that list would benefit.
- Mikhail_K 6y agoIt also leads the readers down the dead-end. Python is incapable of parallelism, the only way to badly emulate it is to launch several runtimes. Writing any scientific software in a badly design language when better alternatives exist is wasted time and effort.
- haihaibye 6y agoIt would be nice to have a faster, parallel Python but the best scientific ecosystems are Python and R. For most scientists, programming productivity matters the most, and plenty of programs are embarassingly parallel. For instance it's no trouble at all just launching a single threaded Python program once per sequencing sample, and it plays nicely with the supercomputer queuing system.
- Mikhail_K 6y ago> the best scientific ecosystems are Python and R. That is debatable