5 ms·
There is some practical relevance to software development here. One shouldn't expose sequential IDs (a.k.a. serial numbers) to the public for anything non-publ
by sparkman55 13y ago
There is some practical relevance to software development here. One shouldn't expose sequential IDs (a.k.a. serial numbers) to the public for anything non-public.
I see this Hacker News post has a numerical ID in the URL, for example; I can estimate the size of Hacker News given enough of these numbers... More directly, I can modify that numerical ID to crawl Hacker News.
Many sites do this; it's generally better to generate a (random or hashed or generated from a natural key) 'slug' to use as the key instead. For example, Amazon generates a unique, non-sequential, 10-digit alphanumeric string for each item in their catalog.
- pavel_lishin 13y agoThe flipside is that you can give off the impression of having a large user base/product catalog/etc if you number things sequentially... but start at a large non-round number.
- fecak 13y agoSome businesses do this with invoice numbers to give the false impression of being more successful (more invoices should equal more sales and customers) than it actually is.
- michaeleloftis 13y agoI did this myself with personal checks. I asked my bank to give me personal checks starting at some moderately sized number so that I wouldn't be giving anyone check #0000001.
- sparkman55 13y agoIt seems like one could use the same technique to estimate the initial (lowest-observable) serial number... From the article: If starting with an initial gap between 0 and the lowest sample (sample minimum), the average gap between samples is (m - k)/k; the -k being because the samples themselves are not counted in computing the gap between samples. Perhaps someone with a better grasp on the math can confirm that this makes 'obfuscating size by starting with a higher serial number' an ineffective mechanism?
- pavel_lishin 13y agoI don't have the math (or time (or ambition!)) to confirm or deny this, but I think that it's not really meant to stand up to serious scrutiny - just enough for someone to glance at an invoice or a URL or a receipt and being more impressed/less wary. (See fecak's reply to my comment.)
- gweinberg 13y agoYes. If you're only looking at the gaps between the numbers, adding a constant offset to the serial numbers would have no effect on the estimate. On the other hand, if instead of ordering them sequentually I roll a die and add the number of spots to the previous serial number, I think I can trick you into thinking I have three times as many tanks as I actually do. In fact, I feel quite confident of it.
- Someone 13y agoIf I find sufficiently many of your tanks, the distribution of the differences in serial numbers would start showing that we aren't talking about a random sample from 1…n For example, having seen 250 IDs in the 1…1000 range and 200 in the 1001…2000 range, the next ID in the 1…2000 range I see should fall in the 1…1000 range with probability 750/(750 + 800) ~= 0.48 in the 'normal' case, and around 36/(36 + 86) ~= 0.30 with your method of doling out IDs. And I think it would be a factor of 3.5 (the expected number of eyes on a throw of a die). That's why I expect your method to dole out 286 out of every 1000 IDs. But it would require me from checking for this, depend on finding such discrepancies in samples, and increase the variance on my estimates for a given sample size.
- cmdkeen 13y agoThe military do this as well. The British SAS when first formed was designation "L Detachment Special Air Service" as part of a disinformation campaign to persuade the Germans there were paratroopers, and at least 11 other detachments of the same size, in North Africa. So there are ways to game the statistical analysis.
- frik 13y agoFor example Facebook exposed the internal user-id, Mark Zuckerberg: http://www.facebook.com/profile.php?id=4 http://www.facebook.com/profile.php?id=4 Other sites like Pastebin use a sequential ID too, but the convert number to another base (like from base 10 to base 43). It's shorter too. (I haven't found the link to the related stackoverflow discussion)
- plorg 13y agoI joined in 2005, when you still needed to be part of a school which had a Facebook network. At that time the ID for most users was made up of a network number (which was in the 4-digit range by the time my school was added) and then a 5-digit user code. So you could (and still can, given the right info) figure out the first person to join from your school, as well as when your school got a network relative to some others nearby. You can still use this to find some mundane information about Facebook's early userbase. For example, the first non-Harvard schools added were, in order, Stanford, Yale, Cornell, Dartmouth, U.Penn, Columbia, etc., and that the first people added to most of those networks either grew up with or attended the same high school as Zuck.
- jjwiseman 13y agoFrom Joshua Schacter, "Autoincrement considered harmful": https://web.archive.org/web/20131029120620/http://joshua.schachter.org/2007/01/autoincrement.html https://web.archive.org/web/20131029120620/http://joshua.sch... Also see https://news.ycombinator.com/item?id=1818166 https://news.ycombinator.com/item?id=1818166
- elwell 13y agoPHP's uniqid() uses time, so that's usually sufficiently indiscernible.
- mikeash 13y agoNonsequential IDs can also help you avoid inadvertently baking in assumptions about your IDs being small, which are then broken as you scale up and cause sorrow and woe. For example, Twitter uses sequential(ish) ID numbers for tweets. Havoc ensued when they crossed 2^32 tweets, because lots of software out there ended up being written with the assumption that a tweet ID could fit into a 32-bit integer. It was true for some years, and then suddenly it wasn't anymore, and everything broke.
- meritt 13y agoI'd advise against modifying that numerical ID. Recent history shows this is a punishable offense (hacking!) under the lovely CFAA and supported by the incompetent judicial system.
- f0xf1re 13y agoFor books, Amazon's ASIN is just the ISBN-10.
- thrownaway2424 13y agoYes there are a huge amount of sequences out there with no special need of being sequential. Apple store invoices. UPS tracking numbers (sharded by shipper). Lots of similar things that can provide a glimpse into the inner workings of what are often publically traded companies with good reason to not leak such information.