3 ms·
The author is making a straw man argument, though it does not seem to be intentional (edit: I am referring to Rubinstein's original analysis). The author points
by sbucher 15y ago
The author is making a straw man argument, though it does not seem to be intentional (edit: I am referring to Rubinstein's original analysis). The author points out that using data for a single class and single year is not a reliable indicator (second graph). That is uncontested and exactly why ratings are based on 3 years of data and multiple classes where possible. Even this cannot provide a single, accurate percentile value. That is why confidence intervals are used and are rather prominently displayed in all the graphs; e.g. <http://www.nytimes.com/schoolbook/school/656-ps-009-teunis-g-bergen/teachers> http://www.nytimes.com/schoolbook/school/656-ps-009-teunis-g...;
Given the example above, there is one teacher who has is 50th percentile for career math, but the confidence interval (CI) indicates that this may really be anywhere from well below average to well above average. In contrast there is another teacher with a value of 3 and the highest bounds of the CI still place them well below average. Conversely there is a teacher that is 90th pctl and even the lowest bounds of the the confidence interval still makes this a highly effective teacher.
In sum, the data make evident that a few teachers are fairly unambiguously ineffective (at least with regards to test results), a few are unambiguously effective, but for the majority the most that can be said is that they are neither the best nor the worst, but somewhere in the middle.
So if you reserve yourself to making conclusions merited by the data, then there is some value to this (i.e. identifying the extremely ineffective and effective). As it stands now, the teacher with a 3 may very well have tenure and be paid considerably more that the teacher with a 90 (given that pay is a function largely of seniority and education, and nearly all teachers have tenure). That is what this exercise is meant to address.
So you definitely cannot assign a fine-grained single value to every teacher (and the authors analyses speak to that issue), but there does not mean you cannot make any conclusions at all. If folks make conclusions that exceed what is supported by the data that is a fault of the analyst or perhaps a function of poor communication or visualization, not evidence that the data is junk.
That said, the test could be improved (and supposedly are being improved) and the very least the measurements place more focus on the subjects being tested.
And the teacher effectiveness data is only 40% of the new approach to evaluating teachers in NY; the other 60% including peer evaluation. It may be useful if that became part of the public data so a more balanced picture is available; but it is unclear if that data would be public.
- kamens 15y agoYou're right, I did take the route of attacking single-metric incentive systems despite the fact that they're not 100% single-metric. I wanted to get a point across that I think still needs getting across. I'm not arguing against the use of data in helping teachers understand their effectiveness for one second, and I mention in the article that the data becomes more correlated and useful as more years are included. "If folks make conclusions that exceed what is supported by the data that is a fault of the analyst or perhaps a function of poor communication or visualization, not evidence that the data is junk." ...completely agree. If my article seemed to argue that "metrics are dangerous" instead of "metrics are dangerous if you choose to publish them publicly while simultaneously using them as a significant component in compensation calculations," then I missed my intended point.
- sbucher 15y agoI was speaking to Rubinstein's underlying article. Sorry for the ambiguity. I strongly agree with you regarding the importance of data being used to empower relevant stakeholders. This issue has been on my mind a lot lately. I would even take it a step further: the priority should be empowering the individuals closest to the data and then working outward. So first priority is making the data empowering for students, e.g. so they have the access and tools to be more reflective about their individual results, learning practices, strengths, and so forth. Depending on age, parents would be here as well. Then teachers. And last should be using the data to empower administrators or bureaucrats. What we have seen is precisely the opposite - the people at greatest distance from the activities generating the data, i.e. the actual learning activities, are most empowered by it. This is reflective of big data business models in general. And I think that is at the heart of where the use of data in education has gone most astray. And this issue is relevant not only to government, but edutech companies as well. Are they maximizing empowering themselves with the data they are collecting to the detriment of empowering students, etc.? Who owns the insights extracted from the data? Are students free to extract their own data and take it elsewhere? etc. It seems Khan Academy has taken this student-first approach as well and has put it into practice. I would be interested in hearing more about there philosophies, practices, or intents with respect to these other dimensions of educational data.