3 ms·
The "maximum number" is a statistic both of a population and of a sample drawn from said population. As the sample size goes up, the sample statistic will conve
by boxy310 11y ago
The "maximum number" is a statistic both of a population and of a sample drawn from said population. As the sample size goes up, the sample statistic will converge to the population statistic, due to the Law of Large Numbers. Because serial numbers are natural numbers (1, 2, 3, etc.) they are uniformly distributed from 1 through N-hat (the population maximum), and the convergence of the sample statistic to the population statistic will scale inversely to the sample size.
Let's take the example in the order it was given in the article: 117, 232, 122, 173, 167, 12, 168, 204, 4, 229
First step we would project the maximum to be 117 + 117 / 1 - 1 = 233. Or, if we drew 1 sample from the population, we would expect the "maximum" to be the average case -- the population mean. We take a minus 1 because there's a degree of freedom lost in using the maximum twice in the same formula (handwaving a bit, but that's a straightforward enough answer for now).
With the second case, we have 232 + 232 / 2 - 1, or 347. We would intuitively expect two cases to effectively be distributed evenly across the population, or at 1/3 and 2/3. (The statistics is a bit different due to the MAX operator here, but I won't go into that).
With the third case, we have 232 + 232 / 3 -1 = 309. Note that while this statistic is still higher than the actual population statistic, it is still converging due to the additional case that helps us counterbalance the weight of the previous unusually-high sample case we observed.
Does that help?
- pinaceae 11y agoso let's say I am being attacked by the first wave of tanks the Germans have fielded. The wave consists of 10 tanks, but from the first batch of tanks produced. 1,5,6,9,12,15,18,20,21,25 how did the allies double check that their sample was not skewed as I described?
- justinsaccount 11y agoYou'd basically be estimating the size of the first batch at that point. Which is fine, until they capture a tank somewhere else that has a number of 230.
- AcerbicZero 11y agoIt should also be noted that there was not one "Tank" SN used, as each component had its own set of serial numbers to be analyzed. In this specific example they used gearbox SN's along with chassis and engine SN to provide the estimate. These statistical estimations were cross-checked with a broad range of other sources, in an attempt to gather data on everything from the number of tanks produced, to the number of factories in use and a variety of supply chain factors.
- mihaifm 11y agodoes this work for consecutive integers that don't start from 1?
- Vraxx 11y agoIt seems that it would trivially work if you knew the number that you start from, you'd just be shifting the uniform distribution (and subtracting whatever number it started from to get the overall estimate). It does seem to be a harder problem if are given that they do not start at 1, but that they are consecutive integers. I imagine you might do a similar sort of statistical estimate on the beginning of the range as well as the end of the range. This introduces more uncertainty, but I see no reason why this wouldn't work
- boxy310 11y agoThere's uncertainty introduced by the idea that you don't know how big that range is supposed to be. Technically the M/n figure should be (M-L)/n, where L is the sample minimum observed, and the estimated population minimum would be represented by L + (M-L)/n - 2 [1]. [1] two degrees of freedom should be lost due to estimating both M and L, thus subtracting 2.