5 ms·
just use random numbers. This is meant more seriously than it might sound. But in the field of "persistent identifiers", there's a notion that languages change
by hwh 10y ago
just use random numbers.
This is meant more seriously than it might sound. But in the field of "persistent identifiers", there's a notion that languages changes over time (a common example being the word "gay"), so introducing meaning into identification schemes might not be a good idea.
- consto 10y agoAlternative why not have both? Use the meaningful url as most links, but have a randomly generated perma link that will always redirect to the same article
- PavlovsCat 10y agoOr sequential numbers, if you don't need to hide anything from being discoverable that way. Or is there another downside to that? I use "slugs" for category type nodes like "photos" or "blog", but I don't give individual "things" a slug (never liked the idea of putting the title into the URL just for SEO). And of course, everything that has a slug always also has the numerical ID and can be reached both ways. Now that I think of it, I might add a table keeping track what nodes slugs were used for, and in case of 404 either redirect when there is only one option, or display all the nodes that once used that slug if there is more than one. Seems like that would be the next best thing to an actually never changing URL, right?
- chriswarbo 10y ago> Or sequential numbers, if you don't need to hide anything from being discoverable that way. Or is there another downside to that? Sequential numbers either require state, which must be kept in sync (e.g. the current-highest index, or the set of all URLs used so far which you can take the "max" of); or else their discovery process is slow (i.e. keep looking up URLs until you find the highest unavailable one). Random numbers don't need state; you can either make the space really huge, and ignore the potential for collisions; or else you can look up the generated URL to see if it's in use, and only have to perform another check in the unlikely case that it is. Whilst it may be easy to maintain the state for sequential numbers (if, say, they already exist as database IDs), the failure mode of random numbers is better. If something goes wrong with the lookup process for random numbers, e.g. it might miss a collision, we're still protected by the incredibly low chance of collisions in the first place. For sequential numbers, we're guaranteed to hit a collision in this case, i.e. if a glitch prevents the state from getting updated, the next URL will definitely collide with the previous URL.
- PavlovsCat 10y agoYeah, I was thinking of auto-increment database IDs actually :)
- MichaelBurge 10y agoIn Postgres, sequences are guaranteed to be monotonically increasing even in the presence of failed transactions. So retrieving an index using nextval()[2] will never collide[1]. [1] A malicious pedant could create situations where it wraps around or collides. [2] http://www.postgresql.org/docs/current/static/functions-sequence.html http://www.postgresql.org/docs/current/static/functions-sequ...
- chriswarbo 10y ago> In Postgres, sequences are guaranteed to be monotonically increasing even in the presence of failed transactions. Well, I was talking about what happens in case of failure; not the chances of a failure to begin with. If, hypothetically, Postgres forgot to update a sequence number, then a collision would be guaranteed, since that's a property of sequential numbering: all updates are contending for the same "next" number, and must be managed in some way. Random numbers spread across a large range, so work even without management. Even if we ignore the possibility of bugs in Postgres, or cosmic rays flipping bits in RAM (which may have a similar probability to a random ID scheme colliding), etc. that's still only true inside the confines of Postgres. Once data starts working its way through layer after layer of shell scripts, caches, proxies, etc. there's a lot of scope for things to go awry. Plus, since we're talking about longevity, there's no guarantee that we'll still be using Postgres in 20 years' time, or maybe we need to migrate or integrate with some other system, etc. I'm reminded of the "one in a billion" chances that used to be claimed for DNA evidence in court. It is certainly true that the probability of two non-twins having identical DNA is that low, but that doesn't really mean much; even if the probability of identical DNA were zero, that doesn't tell us anything about the collison rate of the sub-set of bases which were used, or the error rate of the sampling process, or the mislabelling rate of the lab, or the corruption rate of the evidence handlers, etc. :)
- oneeyedpigeon 10y ago> never liked the idea of putting the title into the URL just for SEO What about for indicating the likely content of a URL, when all someone has is that URL?
- PavlovsCat 10y agoOh, that's a good point. There's really something to be said for hovering a link and getting more than just random numbers and letters. And I know this as a user, but I never thought about the stuff I make in this light.. thanks!
- justusthane 10y agoBecause other people will also link to it using the "meaningful" URL, and when that URL changes the links will break, and now you've defeated the whole purpose of using URLs that don't change.
- Natanael_L 10y agoLet your sitemap page provide a URL history database (directly accessible from any 404 page too!) for every non-permanent URL. That way you can find old permalink mappings for your links.
- leoc 10y agoAnother reason why meaningful names are a terrible idea is because they're often valuable real estate which people will want to reclaim for something else later, or part of valuable real estate like some institutional website's directory hierarchy which the admins feel a compulsion to keep clean and logical and free of crufty old side-passages from years or decades ago.
- Natanael_L 10y agoI feel like the canonical URL for every page should be an UUID, with a database of which human readable document names maps to which UUID:s (and a history of changes should be provided via some tool in case somebody links using the human readable name and it gets replaced, perhaps via the sitemap page).
- tremon 10y agoI feel like the canonical URL for every page should be an UUID Please, not just a UUID. I'd like to know what to expect before clicking a link. I think nu.nl has found an interesting solution to this problem: every article has a canonical url of the form $section/$articlenumber/$title.html -- However, the "title" part (everything between the last slash and .html) is ignored by the server, and is pretty much free to alter as you like. All these three links resolve to the same article: http://www.nu.nl/buitenland/4263268/ http://www.nu.nl/buitenland/4263268/ http://www.nu.nl/buitenland/4263268/amerikaanse-stad-moet-stoppen-met-rassenscheiding-scholen.html http://www.nu.nl/buitenland/4263268/amerikaanse-stad-moet-st... http://www.nu.nl/buitenland/4263268/judge-orders-cleveland-ms-to-merge-segregated-schools.html http://www.nu.nl/buitenland/4263268/judge-orders-cleveland-m...
- Natanael_L 10y agoI almost felt like adding that. I'm guessing appending a title after a # character would be the most technically simple solution?
- tremon 10y ago
- emodendroket 10y agoWell, if the language has changed so much that the URL is meaningless what are the odds the content will be comprehensible?
- FreeFull 10y agoWhy not store things by date, at least?
- icebraining 10y agoThe creation date isn't always that relevant, and may in fact be confusing if the page is later updated.