3 ms·
The only component of student behavior the predictor will not account for is whatever component of behavioral factors which are uncorrelated with SES, but are c
by v21 16y ago
The only component of student behavior the predictor will not account for is whatever component of behavioral factors which are uncorrelated with SES, but are correlated with the distribution of students between classes. I.e., there will need to be a systematic reason why rich white kids who are more poorly behaved than the average rich white kid go to classroom 101 but not classroom 201.
The root cause is that they share a classroom with other students who are misbehaving. To oversimplify wildly, take 4 demographically matched classes:
1) A class with some disruptive students , with a bad teacher : outcome - bad
2) A class with entirely generally well-behaved students, with a bad teacher : outcome - not so good
3) A class with some disruptive students, with a good teacher: outcome - not so good
4) A class with entirely generally well-behaved students, with a good teacher : outcome - good
How do you distinguish between classes 2 and 3?
- yummyfajitas 16y agoHere are two more or less obvious ways. 1) Average over the set of classes taught by any given teacher, maybe drop best/worst (you know, stats 101 stuff). Simple assumptions: do teacher evaluations once per year, 4 classes per semester (so data from 8 classes to evaluate teacher). Assume average scores of 80, and disruptive students reduce scores by 30, bad teachers reduce scores only by 10 pts. Assume disruptive students uncommon, unlucky teacher gets 1 class with disruptive students. Case 2) scores [80,80,80,50,80,80,80,80], mean=76. Mean drop top/bottom = 80. Case 3) scores [70, 70, 70, 70, 70, 70, 70, 70], mean=70. Mean drop top/bottom = 70. Assume disruptive students are common, unlucky good teacher gets 4 classes with disruptive students, lucky bad teacher gets 3. Case 2) Scores [80, 80, 80, 80, 50, 50, 50, 50], mean 65, MDTB=65. Case 3) Scores [ 70, 70, 70, 70, 70, 40, 40, 40], mean 58.76, MDTB=60. Assuming gaps of 5-10pts are statistically insignificant for one year, you occasionally are unable to distinguish between good and bad teachers in a single year. So every 2-3 years, you to a multi-year review. Now your sample size is up to 24 (from 8), and most likely both the good teacher and bad teacher have both had a lucky and unlucky class or two. 2) Include disciplinary problems in the predictor. Thus, the error is reduced from a class with some disruptive students (happens occasionally) to a class with some students who became disruptive for the first time ever (happens 1/12 as often).
- v21 16y agoNo, my point is that disruptive students will drag down all the scores of the entire class. In addition, you have to assume for your corrections that the level of disruption is consistent across subjects and time - ie a disruptive class might collectively be less disruptive when playing PE. Or that the kids change their behaviour over time. Or that there aren't more complex interactions between disruptive behaviour and different teachers (an example might be - kids might be less disruptive just after lunch, because they've just eaten and are sleepy. If you teach "the bad class" in a after-lunch slot, you'll appear better than your colleagues.) And - my (UK) experience, up until the age of 11, is that we had a single "home room" teacher, who taught the majority of our lessons, regardless of subject. We changed home-room (and hence main teacher) once a year. And we haven't even got into the troublesome part of reducing a student's behaviour to a single number. Or the unfortunate way that useful correlates of underlying behaviour stop being useful correlates once you reward people for meeting them. (I'm sure there's a name for this phenomenon, but I have forgotten it) I support the idea of collecting data. I obviously want to analyse it as rigorously as possible. But there are just too many complex interactions going on when you have 30 people in a room trying to learn for our statistics to produce reliable numbers. At least, that's my intuition. It occurs to me that if you actually collected the data, you could do an ANOVA, and have a reasonable stab at attributing the variation in outcomes to various factors, such as individual students, teachers, subjects, the class they're in, interactions between any combination of the above, and "other". You'll need a lot of data, mind. My guess is <10% of the variance would come down to the teacher alone. But this is just a guess - no doubt people have done this before. They've probably done it multiple times, with different answers depending on the different measures they used, whether they adjusted for socioeconomic factors etc. There's probably review articles summarizing those.
- sethg 16y agoThey’re obvious, but they’re not obviously right. All those assumptions that you mention above have to be tested. And every time you throw in another factor to consider, you need to test on a larger sample size to get a valid and reliable model. Look, I’ve worked in the search-engine biz for over five years, and while I myself haven’t gone beyond stats 101, some of my co-workers have studied this stuff at the grad-school level, and they are constantly arguing with one another about how to measure the “quality” of a search engine, and then, given that measurement, what particular kind of statistical model to use in order to predict, given a query and a bunch of potential results, which result has the highest “quality”. (Note that in the particular subfield of search that we are working with, spam pages and black-hat SEO are not really concerns.) And we (like Google and everyone else in this industry) have the luxury of trolling through millions of clicks’ worth of log files and we can hire semi-skilled labor to train our statistical models. And compared with measuring the “quality” of a teacher in a classroom, measuring the “quality” of a search engine is trivially easy.