Tuesday, August 31, 2010

New report critiquing value-added measures

The Economic Policy Institute has released a new report, "Problems with the use of student test scores to evaluate teachers". The report is authored by a number of big names in the education and assessment field, including Dianne Ravitch, Robert Linn, and Linda Darling-Hammond. Click here for the Answer Sheet write-up about the report, which also includes the Executive Summary.

jd

Sunday, August 22, 2010

Income Inequality and Financial Crises

An article from the New York Times: "Income Inequality and Financial Crises".

New technologies, the rise of speculative capital, the polarization of wealth ("income inequality" sounds a bit gentler) are all correlative, and arise together in the hothouse of capitalism. And speculative capital can have a small stabilizing effect, or a major destabilizing one if the sloshing of speculation gets so intense as to spill over into a systemic collapse.

The NYT article referenced above is interesting in that it observes the correlation in the first place, and goes the liberal next step: Poverty isn't just bad for poor people, having lots of poor people around is bad for rich people at a big, system-wide level.

The article is a bit skimpy on how such a causal relationship works.

(The article refers to the mortgage / housing crisis, and here is how the polarization of wealth contributes to financial crises: companies take on more risk, to maintain profitability, by reaching into increasingly impoverished markets, helped along by new technologies that support spreading risk via electronic trading markets. The underlying poverty is unable to sustain the structure, and then it comes tumbling down.)

jd

Falsely identifying "bad" teachers

Here's a link to a recent study from the U.S. Department of Education on likely problems when using test scores to evaluate teacher (and student) performance (referred to in my previous post):

Error Rates in Measuring Teacher and School Performance Based on Student Test Score Gains

The study, by Peter Schochet and Hanley Chiang at Mathematica Policy Research (which develops schemes for value-added measures for school districts, including the District of Columbia Public Schools, where teacher evaluation came into play on the recent firings), has some interesting findings:
  • "Type I and II error rates for comparing a teacher’s performance to the average are likely to be about 25 percent with three years of data and 35 percent with one year of data." [Type I errors are "false positives" -- you think the hypothesis is true when really it isn't; Type II errors are "false negatives" -- you think the hypothesis is false when it isn't.]
  • "These results strongly support the notion that policymakers must carefully consider system error rates in designing and implementing teacher performance measurement systems based on value- added models, especially when using these estimates to make high-stakes decisions regarding teachers (such as tenure and firing decisions)."
  • And this powerful statement: "Our results are largely driven by findings from the literature and new analyses that more than 90 percent of the variation in student gain scores is due to the variation in student-level factors that are not under the control of the teacher."
To reiterate: The first point means that when using three years worth of student test data, chances are about 1 out of 4 that a teacher would be falsely identified as a "bad" teacher. Bad odds when a career and livelihood are at stake.

jd


Some related links on the report:

Study: Error rates high when student test scores used to evaluate teachers from The Answer Sheet blog (very good blog!)

Rolling Dice: If I roll a “6″ you’re fired! from the School Finance 101 blog.

Thursday, August 19, 2010

L.A. Times: Lies lies and more lies

On Sunday, the L.A. Times began publishing an important series of articles on teacher evaluations (Who’s Teaching L.A.’s Kids?). Important because the Los Angeles Unified School District is the nation’s second largest (Chicago is third); important because it appears in a major newspaper; important because they printed teachers names with the data, upping the ante in attacks on teachers (Duncan came out in support of publishing teacher evaluations); important because it previews what I suspect will be the same kind of arguments that will be used by Huberman’s administration against Chicago teachers.

I am late to the party on this -- I first saw a note of it on the District 299 blog (required reading to keep up with CPS news), then on the Answer Sheet blog, and the L.A. Times piece story showed up on NPR yesterday. I’m slow (actually, ironically I suppose, or sadly, I have been mushing and sorting students by their ISAT scores for an upcoming area observation).

The Who’s Teaching L.A.’s Kids? article looks at local student standardized test scores on a by-teacher basis, using a "value-added" statistical model, and based on that model, identifies teachers as "effective and "good" teachers versus "ineffective" and “bad" teachers.

On reading the article, I was reminded of Mary McCarthy’s famous quip about Lillian Hellman (and unfair I think), "every word she writes is a lie, including 'and' and 'the'." Every word in the L.A. Times article is a lie, including "and" and "the". In this case, not unfair.

The article doesn’t have to be "true" of course -- we are in a propaganda war, after all. The tactic of the Times authors is to dress up some statistics bullshit in a pretty hat, and parade it around as science, ergo truth. Everyone is so wowed by the hat, that they fail to recognize that underneath the hat, it’s just, well, just bullshit. But if enough scientistic magic power words are folded into the story, words like "Rand Corp.", "senior economist and researcher", "reliable data", "objective assessment", "effective", the narrative sweeps along and reaches it’s obvious, stinking conclusion.

Here is the essence of the Times’s method:

The Times used a statistical approach known as value-added analysis, which rates teachers based on their students' progress on standardized tests from year to year. Each student's performance is compared with his or her own in past years, which largely controls for outside influences often blamed for academic failure: poverty, prior learning and other factors. Though controversial among teachers and others, the method has been increasingly embraced by education leaders and policymakers across the country, including the Obama administration.
...
The approach, pioneered by economists in the 1970s, has only recently gained traction in education.
...
Value-added analysis offers a rigorous approach. In essence, a student's past performance on tests is used to project his or her future results. The difference between the prediction and the student's actual performance after a year is the "value" that the teacher added or subtracted.


There have been a number of challenges raised to value-added measures on methodological grounds (in particular, misidentification of “bad teachers -- see Study: Error rates high when student test scores used to evaluate teachers from the Answer Sheet blog. Also from the same blog (which is Really Good by the way), Willingham: Big questions about the LA Times teachers project. I have started a list of links here.)

But I think there is a more fundamental, worldview-type fault with the general approach demonstrated in the LA Times article, the Big Lie that makes everything in the article a lie, even "and" and "the". It’s not merely the concept of "value-added" as a metric, but the overall economic approach to education.

Through the lens of economics, teachers go to the education factory. They work on human widgets. At the end of the day, teachers have hopefully added value to the widgets. Value is added if the widgets score higher on multiple choice tests. The greater the change in test scores, the more value that has been added, and the more productive the teacher is.

Implicit in the economic argument is that the education factory must strive to be as productive as possible (i.e. raise test scores as much as possible). Teachers have a greater effect on students than any other single factor, so education reform should focus of identifying the most productive teachers. School districts must then devise incentives to keep the most productive teachers (hence merit pay). Or researchers need to determine what makes for a productive teacher (e.g., Building a Better Teacher, from the New York Times Magazine last March), and teach that in teacher education programs (hence let Teach for America certify its own teachers), and/or suss that out in teacher recruitment or on the firing line in the first couple of years of teaching (see Malcolm Gladwell’s piece "Most Likely to Succeed: How do we hire when we can’t tell who’s right for the job?").

Most all of the research on this approach -- what gives this approach its academic patina of respectability -- points back to the work of Eric Hanushek, an economist at the Stanford's Hoover Institution. He has been working on quantifying the effect of individual teachers, and trying to isolate the teacher effect in education since the early 1970s, and continues to work on it today. His work, and the work of people around him, is the academic foundation, the theory on which most of the official education rhetoric, from Obama to Duncan to Huberman (from what I can tell anyway), is based.

This economic model is taken a step further with the "value-added" notion. I’m not sure where the concept arose but it is an obvious extension of Hanushek’s work. CPS uses a version from the University of Wisconsin’s Value Added Research Center, as does New York City, Milwaukee and Dallas. The LA Times study referred to above was done by researchers at the Rand Corp. A company called Mathematica Policy Research, Inc. was noted in the Answer Blog as the contractor for the Washington, DC teacher evaluation system, which uses a value-added component (and used in the firing of teachers there recently.)

As noted above, "value-added" can be calculated in different ways, but all approaches are based on standardized test scores. As is the case for Hanushek’s work. Here is Hanushek's (and collaborator Steven Rifkin's) statistical justification for saying standardized test scores mean something:

One fundamental question––do these tests measure skills that are important or valuable? –– appears well answered, as research demonstrates that standardized test scores relate closely to school attainment, earnings, and aggregate economic outcomes (Murnane, Willett,and Levy 1995; Hanushek and Woessmann 2008). The one caveat is that this body of research is based on low-stakes tests that do not affect teachers or schools. The link between test scores and high-stakes tests might be weaker if such tests lead to more narrow teaching, more cheating, and so on. (from Hanushek and Rivkin’s Using Value-Added Measures of Teacher Quality, p. 2; emphasis added)


The economic view of education, as the above indicates, assumes the goal in life is earnings and/or academic attainment. If that assumption and mindset is rejected, then the rationale of standardized tests having any meaning evaporates, and the whole argument collapses.

jd

P.S. I skipped over their important caveat re: that the justifying research assumes that tests are low-stakes, which is not the case today with ISAT, ACT, Scantron, etc. The current cheating controversy in Atlanta speaks to the greater incentive to cheat as stakes get higher. In a more perfect world, tests would be part of a bigger assessment profile, and then they might mean something. In the words of Hanushek and Rivkin themselves, the standardized test score data is suspect.

Sunday, August 15, 2010

Computers. Again.

Here are two almost completely unrelated stories (they do both involve computers though):

The First Church of Robotics: A column by Jaron Lanier (remember "virtual reality" anyone?), about the fetishization (or religification) of computers and robots, criticizing the devaluation of thought that takes place when people talk about "artificial intelligence," and reminding us that we "must instead take responsibility for every task undertaken by a machine and double check every conclusion offered by an algorithm, just as we always look both ways when crossing an intersection, even though the light has turned green."

And a much different piece, Market Data Firm Spots the Tracks of Bizarre Robot Traders: The rise to dominance of speculative capital has only been possible with the electronic infrastructure. This article peeks at one strange corner of the world of trading that takes place entirely within interconnected computer systems. Looking at the trading patterns of the bizarre bots graphed out, I wonder if the bot designers are really just using the bots as pencils, to see what kinds of clever patterns they can sketch on their market canvas?


jd

Friday, August 13, 2010

Report on Ren 2010 and charters

The August, 2010 issue of Catalyst Chicago features a number of articles on Renaissance 2010 and charter schools in Chicago.

Also see the "Many Chicago Charter Schools Run Deficits, Data Shows" article by Sarah Karp, deputy editor of Catalyst Chicago, that appears in the New York Times. The article gives a peek at the finances of charters.

jd

Wednesday, August 11, 2010

57,000 monkeys

Here is a link to a Pretty Neat New York Times article:

In a Video Game, Tackling the Complexities of Protein Folding

Basically, someone turned a software application that looked for ways to fold proteins into a video game, so humans could take a crack at folding strategies. Understanding how long amino acid chains fold into proteins is important to understanding the functions different proteins play. (Or something I like that.)

Two things stand out for me: One, humans saw shortcuts and efficiencies that the automated algorithms missed, and brought creativity to the strategies. Two, the Internet provided the infrastructure to allow 57,000 people to do that, together, meaning 57,000 different minds could bring their perspectives and problem-solving skills to bear on the puzzle-like folding action.

jd