10 ms·
Do I really have to cite an arXiv paper?
- docdeek 9y agoAgree with this: if you take an idea from somewhere, cite it. There’s no cost to you personally or professionally for doing so, and you’re giving credit where it is due. >A large number of seminal works have never been published. The greatest mathematics paper of our lifetimes remains unpublished. Is this a reference to a specific paper?
- zackchase 9y agoAuthor here: Grisha Perelman's proof of the Poincaré conjecture has never been (by him, to my knowledge) submitted to or published in any journal. He decided, as is his right, that he could not care less about the professional community of publishing mathematicians or their protocols. Does not invalidate his achievement. I should note, since I am not and do not expect to be the level of mathematician that Perelman is, I have not actually read his proof. So I defer to other superior mathematicians for this assessment and come by it as hearsay. :)
- docdeek 9y agoThanks - I remember reading about Perelman before but I didn’t realise the proof was unpublished.
- acd10j 9y agoI think the logic that his work is unpublished itself is wrong, he published/posted his work for world to see in arXiv , He has very strong opinions about current scientific publishing status quo, hence did not go through usual route of submitting in journal, He did published/posted on arXiv. Who are we to decide that posting your work in Blog post , arXiv etc does not constitute published.
- qznc 9y agoIt becomes a problem if the people with the money want it. Grant proposals usually require you to list your important and relevant "published" papers. There is actually two aspects about "published". One is archival, so people can expect to access the work decades later (if they have to pay for that is another discussion). The second aspect is peer-review aka quality control. Personally, I once submitted a paper to a workshop. After submission, peer-review, and acceptance the workshop committee decided that they will not publish proceedings. I could have submitted the paper elsewhere, which I find weird. Instead I published it as a techreport. However, it is now unusable for proposals, because a techreport is "not published" even if it is properly archived and went through peer review.
- matwood 9y ago> However, it is now unusable for proposals, because a techreport is "not published" even if it is properly archived and went through peer review. IMO, the definition of 'published' is a huge issue. I have always read published as archived peer reviewed research. Your paper is both, but remains in an not published state which I think is wrong and hinders future research.
- backpropaganda 9y agoHey Zack, are there any intermediate/advanced "mathematics for machine learning" books you'd recommend? I find the classic recommendations are not exhaustive enough to cover the kind of math recent papers have started getting into.
- JustFinishedBSG 9y agoThis is very vague. Can you specify what math you are looking for?
- backpropaganda 9y agoThe kind of math that would be necessary to understand papers like [1][2][3][4]. [1]: https://arxiv.org/abs/1703.04933 https://arxiv.org/abs/1703.04933 [2]: https://arxiv.org/abs/1701.07875 https://arxiv.org/abs/1701.07875 [3]: https://arxiv.org/abs/1706.01350 https://arxiv.org/abs/1706.01350 [4]: https://arxiv.org/abs/1512.04860 https://arxiv.org/abs/1512.04860
- yorwba 9y agoYou should keep in mind that probably none of those machine-learning researchers has studied only math specific to that domain, so their papers are likely to include whatever math they have a background in, plus any new techniques they had to learn to get their results. That said, everything I saw in the papers you linked was linear algebra, calculus or probability theory plus the usual smattering of background notation and set theory. Once you have a solid background in those areas, it is likely more productive to look up the specific concepts mentioned in a paper (such as the Kullback-Leibler divergence or the Bellman equation), because by then you are probably too deep in the woods to find one resource that adequately covers all those different directions.
- dsacco 9y agoThat's mostly linear algebra, probability theory and calculus. You're going to have a difficult time self-studying all of that if you haven't had much exposure to it. Books are probably a less efficient method of learning the mathematics if you have targeted subjects you want to learn about. They're typically suited to introductions and breadth-wise coverage of fields, but once you get higher up, "linear algebra" (for example) can get fuzzy with things like abstract algebra. That means you'll end up with several tome-like books to work through which can be productive, but it'll take a while and you'll need to map the material to the applications you're interested in on your own. It's more efficient to develop a good baseline of understanding about a broad subject area, learn the foundational theorems, then move on to the specific areas you need to learn. This is typically doable if you've developed the requisite mathematical maturity overall and if you have learned the "essentials." Practically speaking: maybe pick up foundation texts like Strang's (linear algebra), Spivak's (calculus) and Ross' (probability theory). You're going to want a solid foundation in analysis before moving on to higher order probability theory, so drill down on that after you do a refresher on the calculus. From there you should attempt to read each paper (even if you struggle a lot), take notes on what confuses you or doesn't make sense, read the prior art on those topics and then come back to it. I don't particularly read machine learning papers often, but I read mathematical cryptographic ones very often (at least once per day I find myself in a new one). It's not typical that I read a research paper introducing a novel primitive or construction where I follow the math immediately on a single pass, and I often come across things I need to read about first. From a thirty thousand foot view the math for both of these subjects is broadly similar in rough topical surface area, so I think this methodology for academic reading is fairly applicable to most subjects that involve a lot of mathematics understanding. Basically: don't approach learning the heavy math with a monolithic, brute-force approach as if you were in university. That's a slog and it's demotivating. Learn the minimum foundation for each area you need, then proceed to more advanced topics as you need them.
- smarx007 9y agohttps://arxiv.org/abs/math/0211159 https://arxiv.org/abs/math/0211159 https://arxiv.org/abs/math/0303109 https://arxiv.org/abs/math/0303109 https://arxiv.org/abs/math/0307245 https://arxiv.org/abs/math/0307245
- vermilingua 9y agoPlease use permanent links, this directs to a tag, and will be invalid as soon as a newer publishing article is posted. http://approximatelycorrect.com/2017/08/01/do-i-have-to-cite-arxiv-paper/ http://approximatelycorrect.com/2017/08/01/do-i-have-to-cite...
- devrandomguy 9y agoPlease set up a redirect to the canonical URL. This is helpful to third party systems that interact with your site, such as search engines and content sharing sites. It also eliminates the issue you raised.
- mintplant 9y agoRedirect what to the canonical URL? The entire tag? Why, just to fix this HN submission? The proper solution here is to shoot an email to hn@ycombinator.com (which I've done).
- devrandomguy 9y agoRedirect the tag to the particular article that it currently references. Only show users the "correct" url that you want them to use in the future (history, bookmarks, sharing, etc). https://support.google.com/webmasters/answer/139066?hl=en https://support.google.com/webmasters/answer/139066?hl=en
- arjie 9y agoBruh, the tag is a page containing a list of articles that are tagged with that tag. Redirecting to the latest post with the tag would defeat the entire point of having that page.
- devrandomguy 9y agoThe original link took me straight to the article, not to a list of articles with that tag. It was a single article, with the URL of a tag. Anyway, I didn't mean to argue. Have a nice evening :)
- deleted 9y ago[deleted]
- CogitoCogito 9y agoI am astounded that people would question citing an arxiv paper. If you get an idea from somewhere else, you cite it. This doesn't even have to do with the archive. I have on multiple occasions read papers where the author wrote a proof and then cited "private communications" with another mathematician as the source. In a business where ideas are ideas, you should always cite the source. Anything else is extremely dishonest. edit: To be clear, I am personally entirely okay with not citing papers containing ideas you were unaware of during your own formulation (though I think if you become aware of it, you should probably point out that it was previously independently discovered by someone else). Your paper may still have merit even if it's idea isn't "new" (especially if the first paper is shit as is often the case). edit2: I personally don't see much wrong with Yoav Goldberg's blog post linked in this blog post. It's refreshing to hear his honest opinions out loud. As a graduate student I always lost a little sanity each time I read a paper with "great ideas", but terrible follow-through (i.e. explanation and proof of those ideas). I personally think that clarity of exposition is at least as (if not more) important than the novelty of an idea. However, you should still cite the sources of your ideas. Feel free to point out the source's flaws, but cite them nonetheless. (Hopefully no more edits...)
- hyperion2010 9y agoIn the academy it would be legitimate grounds for an academic misconduct case.
- teniutza 9y agoWhen I started my PhD I was even told not to cite book, regardless of their footprint in the field. They are/were regarded as some sort of "common knowledge". Many ideas come from such works, but one cannot cite them and must, instead, cite a previous work (which surely got their (base)ideas from the same books...). What I'm saying here is that the "academic code of conduct" is a bit outdated.
- CogitoCogito 9y agoI think this is alright as well. Many things do become "common knowledge". The book probably didn't come up with the original idea/argument anyway. Of course tangential to the citing of novel ideas is the citing of material that helps explain your work. I think it's certainly a good idea to cite a source of common knowledge if that source does an especially good job of presenting that knowledge. As with all things, this is a judgment call. Just try to be honest...
- marchenko 9y agoGood article. We want to incentivize researchers to share useful ideas in a timely manner, and allow others to trace the genealogy of these ideas to stimulate their thinking and avoid dead ends. A healthy citation culture helps to build the shoulders of giants.
- pishpash 9y agoIf you were writing a blog post without stinking academic credentialism concerns, would you cite it? If yes, then cite it.
- vmarsy 9y agoIt's a bit tougher when a paper hasn't been peer reviewed, although not too different from a good paper published in a low quality or unknown conference. I think you should cite any idea you pick from a paper when the said paper has some the following qualities: - novelty - technical correctness - clarity - good experimental evaluation Novelty is the most important: if without the citation your paper looks like the original idea, this is plagiarism. If the paper has big shortcomings that your paper adresses, it is fair to give yourself the credit you deserve of course, but it doesn't harm to cite the other paper, in fact it gives a way to give some sort of peer review: In [1], Foo and al. attempted to explore <subject> but the experiments were inconclusive/the technique sucked compare to state of the art/they didn't explain how they did it... In this paper we did this and that and it gives us awesome results (Said in a nicer way)
- taneq 9y ago> I think you should cite any idea you pick from a paper when the said paper has some the following qualities: [...] I think you should cite any idea you pick from a paper, full stop. That's the whole point of citing: To assign credit where credit is due.
- eveningcoffee 9y ago>“Do I really have to cite arXiv papers?”, they whine. “Come on, they’re not even published!,” they exclaim. Is this really a thing? I understood that today you even city web pages (with a time stamp) and even your dog.
- glup 9y agoYes, this is a thing. The existence of a citation format isn't a blanket endorsement of its use in all ways in a scientific work. For example, in linguistics I could cite a specific usage example from a (non peer-reviewed) webpage. On the other hand I wouldn't want to cite a (non-peer reviewed) webpage that outlines a specific linguistic theory.
- anjc 9y agoConfused by this part: > Yes, of course. Any time that our work follows, copies, or borrows ideas from other people, and when we can reasonably be expected to be aware of this, we ought to cite the related work. > We should not have to cite nonsense. Many reviewers are abusing the system and asking for ridiculous comparison to recently-posted preprint papers. Bald-faced flag-planting should not be rewarded. And we should not be faulted by reviewers for failing to compare against 2-week old algorithms that may or may not work. So what position is the author advocating? Citing or not citing? edited
- detaro 9y agoInclude the next sentence for the first quote: Yes, of course. Any time that our work follows, copies, or borrows ideas from other people, and when we can reasonably be expected to be aware of this, we ought to cite the related work. Something being on arXiv or in a blog post or … doesn't excuse not citing it if it influenced your work. It's important to document where your ideas and data come from, both to give credit to the author and to allow others to evaluate what you base your claims on for themselves. The second quote is about stuff that didn't influence your work. While you are expected to keep up with and document related developments, forcing authors to constantly update references to new, not yet properly evaluated work just because it makes some related claim doesn't make sense.
- anjc 9y ago> The second quote is about stuff that didn't influence your work. While you are expected to keep up with and document related developments, forcing authors to constantly update references to new, not yet properly evaluated work just because it makes some related claim doesn't make sense. It doesn't say that though, it just says that you shouldn't have to cite nonsense on arXiv. This implies that you can read a paper, implement something similar, and afterwards decide that the paper was nonsense, didn't influence your work, and shouldn't be cited.
- detaro 9y ago> Many reviewers are abusing the system and asking for ridiculous comparison to recently-posted preprint papers. Bald-faced flag-planting should not be rewarded. And we should not be faulted by reviewers for failing to compare against 2-week old algorithms that may or may not work. The context seems pretty clear to me. Your example clearly is covered by the first case: if it influenced your work, you cite it, even if it is "nonsense". You can't just "decide" something didn't have influence if it had. (These rules do not prevent cheating, they are guides for people acting ethically) The second rule is to prevent the opposite case: You shouldn't be forced to create the impression your work is based on or just a mere repeat of someone else's "who had the idea first" when they have no good claim to that, or inferior to something that hasn't been shown to be actually better.
- a_e_k 9y agoI tend to view arXiv as mainly an aggregated repository of documents that would ordinarily be tech reports. It is accepted practice to cite tech reports when the paper author is aware of them and they are relevant. I don't really see how an arXiv paper should be any different.
- SeanMacConMara 9y agoHow is "flag-planting" even relevant to this ? You only cite what you use and you should not be using vapid hollow flag-planting sources from _anywhere_. Someone needs a fresher course on academia 101. Is it the reviewers ?
- zackchase 9y agoFlag-planting is relevant because the practice of flag-planting (and other abuses of the pre-print) is what has prompted some of this over-reaction.
- qznc 9y agoThere is a lot to be fixed about reviewers, but I have no clue how to do that.
- mattkrause 9y agoThe issue is that some people (appear to be) submitting very preliminary, and arguably low quality work to arXiv to stake a claim to some area. Once that work is "out there", people who were making a more serious effort to do things more carefully are obligated to cite the original work, presumably as part of the related work/background/etc. I can see how this would be maddening, particularly if you started before the flag-planting paper was even written.
- lmm 9y agoThere are two kinds of "cite" here. Citing a non-academic source is different from citing a (published) paper; you should cite anything that precedes your work in the second sense, whereas the first is only obligatory for work that you actually took something from. If you work on calculus you're obliged to cite Leibniz even if you didn't read him, but you are not obliged to cite Newton's unpublished work unless you read it. Unpublihed arXiv papers fall in the latter category.
- nl 9y agoUnpublihed arXiv papers fall in the latter category. In CS that almost certainly isn't true. I'm most familiar with the NLP field, but there, if you have some kind of embedding of your words/tokens/sentences/something you cite https://arxiv.org/pdf/1301.3781.pdf https://arxiv.org/pdf/1301.3781.pdf (Word2Vec, Mikolov). That paper says there is a follow up paper published at NIPS2013, but I don't think I've ever seen that published. The field just moves too fast to wait for conferences anymore.
- mathperson 9y agothis is just plain wrong.
- sgt101 9y agoThere is a massive disjunct between citation in the humanities and citation in science. People seem to have forgotten this. Ideas are two a penny, they literally do not matter at all they should not be cited. The person who introduces a concept to science deserves no credit whatsoever. What deserves credit is the provision of evidence or proof.
- lou1306 9y agoWhat? Leibniz/Newton introduced the concept of calculus; shouldn't we credit them with this (amazing) concept just because it wasn't formalized very well until Weierstrass came along?
- analog31 9y agoIn my view, Leibniz / Newton did enough heavy lifting to deserve credit for developing the field. This is way different than just blurting out a "concept" without even knowing if it's workable. I'm in agreement that ideas are a dime a dozen. Sure, it's necessary to cite the first known mention of an idea, and there are situations where the first mention of an idea is important such as in the patent system. "Planting" happens in my world all the time. Unfortunately, managers give a lot more importance to "ideas" than they are really worth, because they over-value their own interventions in general. Somebody will blurt out an idea in a meeting, wait until someone else has developed it, and then rush in to take credit. If a manager does this, it's a blow to morale. I have my own rule of thumb, which is "show your work" from math class. Just writing down the answer doesn't get you full credit. A couple of historical examples: The ancient Greeks are credited with the atomic theory, but they had no concept of even turning it into a serious hypothesis. Lots of ideas are anticipated in science fiction, but do those authors really deserve credit?
- sgt101 9y agoI agree, both Newton and Leibniz worked out chunks of method and demonstrated their utility. Interestingly Newton's alchemical approach to publication is rather similar to an arXivists... bits and snippits! In the end we reference the formal publication (if we are developing fundamental changes to calculus which would be f*ing impressive... or writing about science!)
- kutkloon7 9y agoMy masters thesis is based on a single arXiv paper. I didn't know that arXiv had this questionable reputation, but it sure explains a lot. The project is to make an FPGA implementation of a technique presented in an arXiv paper. The paper had some big gaps, so I had to spend a lot of time researching the technique, and I had very little time to spend on the actual implementation.
- skgoa 9y agoarXiv is just a repository for papers to make them available before they are published/peer-reviewed. Just because a paper is on arXiv does not make it bad, but it doesn't mean that it is high quality either. Typically the people in the field will know which papers are important.
- kutkloon7 9y agoIt's not bad, there are just some gaps. I expect a follow-up paper which explains everything thoroughly. I feel like my masters thesis should have started after that paper.
- chriskanan 9y ago"If you built on it, cite it", is necessary but does not suffice. Most papers have a related work section that describes all work similar to your paper. In hot areas, such as deep learning, some of the related work may have been done in parallel with your work, so it did not inform it. If this related work is uncited on arxiv, you find it when you are about done, and it had no influence on your work, do you cite? Reviewers sometimes demand this. I've been told a rule of thumb is that if a related unrefereed arxiv paper has been cited six or more times, with the justification being that this means it is somewhat well known once it has some citations.
- greeneggs 9y agoOf course you cite it. You want to help the reader find the related work. It doesn't matter whether it "had influence" on your work (a fuzzy criterion). If you aren't happy that you are doing the same work as others, then find a more original problem. It definitely does not matter how many citations the other paper has. The point isn't to avoid getting caught, it is to inform the reader. Your citation is more useful the less well known the cited paper is.
- nthcolumn 9y agoI must be missing some nuance of the argument here. If it is nonsense then why is it in his paper? If you are citing poor quality sources then that tells us something about your paper and to not do so would be dishonest. We're not seriously saying that you should be trawling for similar ideas to your own and then citing them, just where you have used other work you must give credit.
- sytelus 9y agoThe main problem is giant publication latency. Typically there is at least 6 months delay by the time you submit paper, it gets reviewed and then actually published with proper DOI etc. These days 6 months is loooong time. I wish there was some way to generate all relevant bib information as soon as paper gets accepted which then can be added on arxiv immediately. This would allow folks to distinguish between peer reviewed papers vs those which are submitted only for flag planting.
- lotsoflumens 9y agoCiting an arXiv paper should take precedence over citing anything else. If you really believe in science then ONLY non-paywall papers should be cited.
- mindcrime 9y agoI don't see how this is even a question. A paper "published" to arXiv is published, in the more general sense of the word. Just because it isn't "journal published" doesn't change anything.
- glup 9y agoPublishing in a (reputable) journal generally means some degree of peer review. While noisy, this process generally means that really outrageous methodological errors or theoretical claims get weeded out. For a good paper, it means that other researchers have pressed them on specific aspects of the work, which often produces stronger work (new methods, better baselines, clearer argumentation, clearer math). I adore arXiv but still believe it's a preprint. In my field (cognitive science) it would be great if we had more methods to sidestep Elsevier and the other commercial publishers and have an open stack with rigorous peer review (PLoS being the main way currently).
- cwyers 9y ago> While noisy, this process generally means that really outrageous methodological errors or theoretical claims get weeded out. There is a fair amount of evidence that this isn't true. In general, most statistics in scientific research aren't done by statisticians, and there are whole classes of methodological errors that are regularly not caught because the "peers" have the same lack of statistical education as the people whose papers they are reviewing.
- mindcrime 9y agoWhile noisy, this process generally means that really outrageous methodological errors or theoretical claims get weeded out. For a good paper, it means that other researchers have pressed them on specific aspects of the work, which often produces stronger work (new methods, better baselines, clearer argumentation, clearer math). I see that as all true, but irrelevant in this context. If you source material from a pre-print on arXiv, then you should cite it. Seems totally obvious to me. Of course you would prefer the final, published paper if it's available. But that wasn't the question at hand. And even with all that said... I would argue that in some fields, (cs / ml / etc.) we're getting close to a point where arXiv itself is become almost a parallel publishing mechanism where people cite/publish completely within the arXiv realm, with less regard for "traditional" journals and what-not in general. Especially when you factor in papers from researchers who come from industry, as opposed to academia, and care less about some of the normal trappings of academic publishing. I adore arXiv but still believe it's a preprint. Of course it's a pre-print. I didn't contend otherwise. I'm just saying that, from my perspective, it's obvious that you should cite a pre-print if it's relevant. I will allow though, that norms probably vary from field to field, and as a non-academic, my take is likely different from, say, somebody who is deeply immersed in academia, pursuing tenure, etc.
- al2o3cr 9y agoSeems like this is a case of "whatever you measure will be gamed": counting citations is an important part of how academics are evaluated at work, so we get flag-planting behavior to maximize that metric with minimal effort. There's a similar issue in journal publishing: counting "published works" without regard for where leads to journals that will publish literally anything for cash.
- dekhn 9y agoIn the future, the question will be "Do I really have to cite <paper in a closed journal>"?
- avmich 9y agoWhat if somebody independently came to similar or worse results and only then read about the research - with possibly more results - made before that? Given the amount of information available, it could often be the case of independent research into something which is known and available for some time. If one only learns about similar - and possibly greater - results after making one's own, and wants to talk about the work done - should one cite other, possibly earlier, works?
- erik998 9y agoYou definitely should cite... What if they are wrong... You can always blame the citation... Polywater is the perfect example of why you should citate. Just because someone publishes something does not mean its correct, scientific, or proven. https://en.wikipedia.org/wiki/Polywater https://en.wikipedia.org/wiki/Polywater http://science.sciencemag.org/content/167/3926/1715?sid=8b4eadf1-7198-4b31-b0fe-e0b27d28b8cf http://science.sciencemag.org/content/167/3926/1715?sid=8b4e... Science takes time. Write a good paper and cite, cite, cite! If your citations are ever demonstrated to be incorrect or fraudulent other researchers can continue work to disprove the citations and work on correcting the errors.