4 ms·
Can someone give me an example on how to properly use the central limit theorem in the real world? For example with stuff that are commonly assumed to be bell-
by fyp 7y ago
Can someone give me an example on how to properly use the central limit theorem in the real world?
For example with stuff that are commonly assumed to be bell-curved like test scores, IQ, etc. What are the iid variables being averaged? Each test question?
- throwlaplace 7y agothere's a distinction between random variables that are assumed to actually be normal and RVs that we conceptually apply the clt to and then treat as normal. in the case of test scores it could be either. for an individual test it's possible the distribution is roughly normal actually. but if you, for example, look at subsets of all SAT scores and take their means (the mean of the subset) then those averages will be Bell shaped because of the clt. Just to directly answer your question: the iid random variable is the test score or the iq.
- markkvdb 7y agoSuppose you have a school with 30 groups of students. All students are randomly assigned to the groups (so independent of their skill at IQ tests). The distribution of measured IQ within classes can be any distribution. But let’s take the 30 averages of the groups. The mean of these averages of groups (asymptotically) follows the normal distribution according to the CLT.
- throwlaplace 7y agonot the mean of the means - just the means themselves. The distribution of the original population of iqs in the school is what you're sampling. Each group is a sample. The distribution of the sample mean approaches a normal according to clt as the sample sizes increase.
- deleted 7y ago[deleted]
- raegis 7y agoHe's the example I have my students do in class. Roll a die repeatedly and tally the results. You'll get a (roughly) uniform distribution of 1s, 2s, 3s, 4s, 5s, and 6s. Now to illustrate the CLT, you roll a die 50 times, and average the result. AND your 300 classmates do the same. If you tally the 301 averages, the distribution of the averages will not be uniform but bell-shaped, with average (approximately) 3.5. The CLT says (roughly) the distribution of the averages will be approximately normal, regardless of the original distribution.
- ganzuul 7y agoShort and concise, as things should be. Thank you. :)
- laichzeit0 7y agoHere’s something that’s bothered me for a while. Why is it necessary to roll it 50 times? Why can’t I roll it just once? Does each X_i have to be a sample greater than 1 and if so is there some kind of minimum required for each random sample? Does CLT still hold if each random sample’s size is 1?
- zosima 7y agoAs the number of samples approach infinity the mean of all the iid samples will become normally distributed. So it will neither hold strictly for n=50 or n=1, but n=50 may be a better approximation of infinity :)
- laichzeit0 7y agoAs I understood it each X_i is a random sample and the n refers to X_n, that is, n random samples. But what is the size of each sample? In the parent it was 50. The dice is rolled 50 times by each person i, and the outcomes are given by the random variable X_i. Also, there are 300 people rolling a dice, so n=300, thus 300 random variables each composed of 50 dice roll outcomes. I’m curious why each X_i has to be a sample of size 50 and can we just have each person roll the dice 1 time? Maybe we need 50*300 people to roll the dice one time now? Does the CLT work when each X_i is a random sample of size 1?
- roenxi 7y agoThe CLT can be used to justify modelling an effect with a normal distribution and raises questions about model assumptions when normal distributions don't appear where they should. The Normal Distribution is special because it is the highest-entropy distribution that is completely characterised by mean and standard deviation [0]. So what the CLT is saying is that, surprisingly, sums of i.i.d. random variables lose information from the individual variables quite quickly but reliably retain a little data about mean and second moment. Initially a given X_i has all sorts of information associated with it (higher order moments, other distribution characteristics, etc) that disappears as many X_i are summed together. All that is left is information about a mean and variance. So say I have a situation where a large number of i.i.d. variables are going to be summed together. The CLT tells me that summing N of these variables together is going to be similar to summing together an equivalent number of normally distributed variables (!!). This is because the sum variable is equal in value to n * \bar{x}, but the CLT imposes a distribution on \frac{\bar{x}}{n}. This justifies why the normal turns up everywhere in practical measurements. A lot of measurements (say, number of people at the beach) are probably really measures of a sum of random variables (maybe there is some non-normal variables that captures the chance a given person goes to the beach). So if the number of people at the beach turns out to be normally distributed (maybe I measure it each day for a few weeks) it isn't shocking. If the total number isn't normally distributed then that implies that there is no i.i.d. variable representing probability a given individual goes to the beach (eg, high correlation between the individual variables). I didn't quite fail any of the statistics courses I've ever done, YMMV. [0] https://en.wikipedia.org/wiki/Normal_distribution#Maximum_entropy https://en.wikipedia.org/wiki/Normal_distribution#Maximum_en...