Interesting article, though I think it incorrectly leaves
the reader thinking that there is some interesting
informating hidden in the average spacing of the numbers.
In fact, all you need to know is that maximum observation
and the number of observations. Once you simplify the average
spacing goes away.
If M is the maximum serial number of N is the total number of
observations, using the formula in the post:
M + (avg. spacing) = M + M / N - 1 = (N + 1) / N * M
To me that gives a more clear picture of what the unbiased
estimator is doing: inflate the maximum value by a factor that
limits towards one as the sample size grows.
If you just assume that the sample mean = the population mean, then you get the right answer, at least for this example. I don't see why the article fools around with the maximum at all - isn't the maximum a much more noisy statistic than the mean?
The range matters – had they found 10 serial numbers between 100000 and 101000, would the mean still be a meaningful estimate of the production rate? In this case, the author just tacitly assumes the minimum to be zero.
To be the devils advocate: what you say is true if you know the distribution. If spacing looks weird (e.g. clustered) it might indicate that the number is, for example a pairing of model and serial numbers, etc.
Comments
Interesting article, though I think it incorrectly leaves the reader thinking that there is some interesting informating hidden in the average spacing of the numbers. In fact, all you need to know is that maximum observation and the number of observations. Once you simplify the average spacing goes away.
If M is the maximum serial number of N is the total number of observations, using the formula in the post:
To me that gives a more clear picture of what the unbiased estimator is doing: inflate the maximum value by a factor that limits towards one as the sample size grows.If you just assume that the sample mean = the population mean, then you get the right answer, at least for this example. I don't see why the article fools around with the maximum at all - isn't the maximum a much more noisy statistic than the mean?
The range matters – had they found 10 serial numbers between 100000 and 101000, would the mean still be a meaningful estimate of the production rate? In this case, the author just tacitly assumes the minimum to be zero.
To be the devils advocate: what you say is true if you know the distribution. If spacing looks weird (e.g. clustered) it might indicate that the number is, for example a pairing of model and serial numbers, etc.
Distribution of manufacturing date or distribution of rate of tank capture?
Or does it make a difference?