4 ms·
> This is a weird concern to raise -- the idea that EF's data isn't representative because it's pulled from a sample of citizens extraordinarily interested in E
by mikk14 7y ago
> This is a weird concern to raise -- the idea that EF's data isn't representative because it's pulled from a sample of citizens extraordinarily interested in English would strongly imply that the true proficiency level in every country is much lower than reported, and to a lesser extent that the gaps between countries are probably wider than reported.
No, if your samples aren't representative then any conclusion you draw can be flawed for thousands of different reasons and in all possible directions. Two examples.
1) Let's say that the top 1% of English speakers of countries A and B took the test. You say "the gaps then would be wider than reported". Wrong, even if in this 1% sample country A scores the highest, it might very well have lower overall proficiency. If proficiency is normally distributed, but has a much greater variance in country A than country B, the maximum of A will be higher than of B, even with a lower average.
2) Now you take the tests at time t and t+1. However, at time t the test was less popular, so only 0.5% of the top speakers took the test, while at t+1 1% of the top speakers took the test. You would conclude that a lower score a time t+1 implies that the skill got worse, while it might have very well improved by a lot.
(In fact, I'm pretty sure that's what happened. I never heard of EF before 2017, so I couldn't have taken the test before then)
- thaumasiotes 7y ago> 2) Now you take the tests at time t and t+1. However, at time t the test was less popular, so only 0.5% of the top speakers took the test, while at t+1 1% of the top speakers took the test. You would conclude that a lower score a time t+1 implies that the skill got worse, while it might have very well improved by a lot. I agree with this in full. I don't see what it's responding to in my comment, though. > 1) Let's say that the top 1% of English speakers of countries A and B took the test. You say "the gaps then would be wider than reported". Wrong, even if in this 1% sample country A scores the highest, it might very well have lower overall proficiency. If proficiency is normally distributed, but has a much greater variance in country A than country B, the maximum of A will be higher than of B, even with a lower average. I'm not following you here. I claimed that when the test pool for each country is strongly self-selected for interest in English, the gap between the countries' average proficiency levels (in reality) is very likely to be larger than the gap between those countries' average test scores. The test scores are suffering from restriction of range. An example in which country A has higher test scores than B at the same time it has lower proficiency looks like an example of that phenomenon, not a counterexample. If A and B have the same mean proficiency and different variance, and selection into the test pool is driven only by proficiency, then the country with higher variance will test better under this system while having the same mean proficiency, that's true.