3 ms·
There is probably something to this model. Notice first that the constants are suggestive. There are 52 working weeks in the year, and a slightly more than 250
by pash 12y ago
There is probably something to this model. Notice first that the constants are suggestive. There are 52 working weeks in the year, and a slightly more than 250 working days, and the author finds a good fit for
sick_days = -52 * log(x) + 236 .
So the fitted curve can interpreted as saying,
sick_days = weeks * sick_days_per_week(x) + healthy_days + noise ,
or somewhat more straightforwardly,
days = healthy_days + sick_days(x) + noise .
Since x is the rank (index) of an employee by the number of sick days taken, we have that
sick_days_per_week = log(employee_index) .
So it seems that we should guess that sick_days_per_week is log-normally distributed across employees. That's the same as saying that employees take sick days at a rate
r = mu + sigma * Z ,
where mu and sigma are the mean and standard deviation of sick_days_per_week and Z is a standard normal.
This checks out intuitively, since this model comports with one in which the number of employee-sick days in any timeframe is independent of the number taken in previous periods. And the author says his sick-day policy is set up to incentivize taking sick days only if you're actually sick, not because you've accumulated them, etc. So the data look like what we would expect if the policies are working as they should. (Notice also that this model eliminates the need to control for employee tenure.)
So hopefully the author will run the data and see whether Z passes statistical tests for standard normality, then tell us what mu and sigma are. We might have a nice model for employee sick days and some solid empirical results to go along with it!
- contravariant 12y agoFrom his claim that sick_days = -52 ln(x) + 236 It would be more straightforward to conclude that the number of sick days is exponentially distributed. To see this first note that the employee index "x" is equal to N(1 - p), where N is the number of employees and p is proportion of employees with less sick days. So if we invert the formula we can calculate the proportion of employees with less than a certain number of sick days, which is basically the cummulative probability distribution of the number of sick days. Now let's try to invert the formula. First we simplify a bit: -52 ln(x) + 236 = -52 ln(N(1-p)) + 236 = -52 ln(1-p) + 236 - 52 ln(N) Now if we fill in N = 86, we see that 52 ln(N) is 231.626 which looks suspiciously like 236, so let's assume that they're the same for now. This gives us: sick_days = -52 ln(1-p) Or in other words p = 1 - exp(-52*sick_days) Which is the cumulative probability distribution for an exponentially distributed random variable with λ=1/52. So it seems that the number of sick days is approximately exponentially distributed with an average of about 1 week. It's also satisfying to see that the constant "236" almost completely follows from this assumption. There was still a minor difference, but that can be removed by assuming that the number of sick_days is at least 4.37.
- pash 12y agoGood catch that the additive constant is approximately equal to 52 * ln(86). I like your model. It makes sense, it's simpler, and most importantly it's different enough from mine that it gives us a decent way to test which model reflects reality better: fit data drawn from different numbers of employees' sick-day history. My model is independent of the number of employees in the set and yours isn't, so that should tell us something. For those of you following along, note that it's often rather difficult to tell what the "real" model is, even assuming that there is one and that we've somehow managed to discover it. Like many other pairs of distributions, the exponential and log-normal distributions are quite similar, both in how they fit data and in the intuition behind them. There are choices of parameters that make an exponential and log-normal distribution look about the same, so without good theoretical justifications for which parameters to choose, there's little reason to prefer one over the other. Each can be thought of as giving a time to failure (here, the time between sick days taken). The log-normal has two parameters and the exponential has only one, so if neither model is "correct", it's likely that we could find a choice of parameters that makes the log-normal fit even if the exponential doesn't. The two distributions differ most in the tails, but we would need gobs of data to see how the tail probabilities work out. If other employers with some statistical expertise want to weigh in with their own employee datasets, please do! Edit: By the way, I think this back-and-forth is a good example of the benefits of dialectical reasoning (that is, a conversational process of trying to hone in on the truth, with a real interlocutor to converse with). I came up with my initial model only because the two constants seemed curiously coincidental with the number of weeks and working days in the year; contravariant's model throws one of those away, which is fine, but I likely would not have proposed my model in the first place if only one constant looked suggestive—a singular "52" does not really get the modeling juices flowing. And I suspect thst contravariant would not have thought up his model if he hadn't seen mine. As someone who mostly works alone, I wish I had more of these sorts of dialogues about the problems I work on. Edit 2: On re-reading my grandparent post, which I can no longer edit, I just noticed that I made a very fundamental mistake that means my model is certainly incorrect. To see my error, start with my third equation, which is obviously correct (except that there can't be any noise). Then plug in sick_days as defined in the first equation, which comes directly from the blog post. Now try to derive my second equation. ... contravariant did not make this mistake.