4 ms·
The reason of the apparent paradox: a) in this case the model is a mixed model b) second the variable are nominal so you have to select one of the pseudo R^2 mo
by justk 2y ago
The reason of the apparent paradox:
a) in this case the model is a mixed model
b) second the variable are nominal so you have to select one of the pseudo R^2 models.
For more information:
(1) Pseudo R-squared: https://en.wikipedia.org/wiki/Pseudo-R-squared https://en.wikipedia.org/wiki/Pseudo-R-squared
(2) R squared for mixed models – the easy way https://ecologyforacrowdedplanet.wordpress.com/2013/08/27/r-squared-in-mixed-models-the-easy-way/ https://ecologyforacrowdedplanet.wordpress.com/2013/08/27/r-...
c) The R^2 used with a linear model requires a constant term, in this case the constant term or bias explains a lot about preferences (almost 50/50) so there is less information available for the slope term.
Hope this helps.
- vcdimension 2y ago@justk What you're talking about might make sense if there were more independent variables to consider, but in this case there's only one, state. So in fact you could say that there are two conditional linear models in the example; one for the first state (state=0), and one for the second (state=1). The model does the best job with the information available (state).
- justk 2y agoSorry, I edited my post several times and finally choose a short form with links other sources. If you fix state=1 then there are no more random variables so the R^2 doesn't have any meaning. Just for fun, what the model should predict for state = 0.5?, that corresponds to a person that is 50% in the red state and 50% in the blue state, I think a mixed model is appropriated here when the state variable is discrete, so that each value of the state variable represents a different part of the population, the other model should be used when people move a lot and change frequently the state where they vote in, but in that case you should have to consider the fluctuations in the total population in each state at the time of voting.
- vcdimension 2y ago@justk The R^2 value of 0.01 calculated on that webpage uses both states, not just one: the variance of the predicted values across both states is 0.55^3+0.45^3 - (0.55^2+0.45^2) ≃ 0.497 ≃ 0.5 I don't think it makes sense to use a mixed model in this case since the variance is the same for each state. A mixed model is used when the observations have some structured heteroskedasticity, i.e. different variances for different values of the independent variables.
- deleted 2y ago[deleted]
- fjkdlsjflkds 2y ago> The R^2 used with a linear model requires a constant term, in this case the constant term or bias explains a lot about preferences (almost 50/50) so there is less information available for the slope term. This explains the paradox, basically. When you take the null model "preference = 50%" (i.e. intercept-only model), there simply isn't much residual variance left for the linear model to explain. That's why you get an R^2 = 1 if you use the "R^2 = rho(state, preference)^2" formula (you are ignoring the role of the intercept in explaining most of the variance, and exploiting the translation-invariance of the Pearson correlation) vs. you getting an R^2 = 0.01 when you use the (more correct) "R^2 = explained variance / total variance" formula. TL;DR: It makes sense to get a very low R^2 when it is the intercept and not the predictor that is explaining most of the variance.
- kgwgk 2y ago> there simply isn't much residual variance left for the linear model to explain. I'd say that there is still quite a lot of residual variance to explain. You need a baseline - the worst choice would be to predict 0 (or 1) and the mean squared error would be 0.5. Using 0.5 as baseline halves the mean squared error to 0.25.
- fjkdlsjflkds 2y agoThis, of course, will depend on how you code your variables, but if you try to fit a null, intercept-only and predictor-only model, you get this as residual variance: > data <- data.frame(state = c(0, 1), pref = c(0.45, 0.55)) > sum(residuals(lm(pref ~ 0, data = data))^2) # null model [1] 0.505 > sum(residuals(lm(pref ~ 1, data = data))^2) # intercept-only model [1] 0.005 > sum(residuals(lm(pref ~ state + 0, data = data))^2) # predictor-only model [1] 0.2025 So, it seems clear that you only get a "perfect" prediction with the full (intercept + predictor) model mostly because of the intercept (which explains (0.505-0.005)/0.505 = 0.99 = 99% of the variance). Thus, it makes sense that the predictor is only explaining the rest (i.e. 1%) of the variance... hence, the R^2 = 0.01
- kgwgk 2y ago