5 ms·
It's really hard to attribute all these posts to the same content when they use different URLs.
by phgn 4y ago
It's really hard to attribute all these posts to the same content when they use different URLs.
- ColinWright 4y agoThese are just the top few in the search engine's rankings. If you search for the URLs you'll see that some of them have been submitted many, many times. For example: * 12 times: http://www.catb.org/jargon/html/story-of-mel.html http://www.catb.org/jargon/html/story-of-mel.html * 6 times : https://www.cs.utah.edu/~elb/folklore/mel.html https://www.cs.utah.edu/~elb/folklore/mel.html * 4 times : http://www.pbm.com//~lindahl/mel.html http://www.pbm.com//~lindahl/mel.html * 3 times : http://www.jargon.net/jargonfile/t/TheStoryofMel.html http://www.jargon.net/jargonfile/t/TheStoryofMel.html It's just that subsequent submissions have not always scored as highly, exactly because people recognise them. Perhaps "classic" should be: * Has scored more than 50 points on some submission[0]; * The URL has been submitted more than 10 times[1]; ======== [0] For some value of "50" [1] For some value of "10"
- cxr 4y agoThis is the tip of a much bigger set of related, often frustrating problems. It's a problem with annotations in a loose sense, for example, which involves anything that relies on URLs as identifiers. The reality is that URLs _are_ identifiers, but it's necessary to recognize that they identify different "printings" (e.g. a work carried by many different "publishers"—The Story of Mel carried by Lindahl is not the same printing as the one carried by Brunvand, for example). There's a lot we're missing out on by abandoning the lessons (and conventions) of legacy print media in the move to the Web. The addition of "accessed on $DATE" in bibliographic citations is an indicator of how we've truly messed things up with this abandonment and by our desire to treat the Web as a sui generis medium unlike anything that came before. Both this and the industry at large has normalized/legitimized (professionalized, even) unhygienic practices in publishing. What's needed is probably something like the way OpenLibrary functions as a registry for works and particular editions, or an open database like graph.global where you can can mint relations linking individual pieces of content—this thing (identified by URL $X) and this other thing (identified by URL $Y) are the "same" thing. However, I think we'd most benefit as a first step from a technological and social re-calibration of the way we interact with URLs entirely: <https://news.ycombinator.com/item?id=29803419 https://news.ycombinator.com/item?id=29803419> (It's telling that the "past" links at the top of all HN submissions point to a search by title, rather than a search by URL. I've personally hit minor hurdles in my own Algolia searches that turn up 0 results, only to eventually realize that it's because the URL I've entered differs from the one that was submitted in their respective "https"/"http" schemes.) The convention I outline in that comment is probably the way to go. If I could refer to <https://apress.com/P. Seibel. Coders at Work: Reflections on the Craft of Programming. Apress, 2009. https://apress.com/P. Seibel. Coders at Work: Reflections on...> then that would help a lot. It wouldn't fix the deduplication/canonicalization problem, exactly, but it would reinforce some good habits and the way we conceptualize the things we're working with, which could lead to this sort of thing becoming more tractable.