The White House asked for some feedback on their proposed "scorecard" for higher education cost and value which is intended "to make it easier for students and their families to identify and choose high-quality, affordable colleges that provide good value." Below are their questions and my (quick, off the top of my head, answering-an-online-survey level of analysis) responses.
What information is absolutely critical in helping students and their families choose a college:
You shouldn't be asking this question here. It's a researchable, empirical question. First, on what basis DO people decide? Then, to what degree do they have the appropriate information to do so?
As someone who studies things like this, I don't think the info presented here as it is here presented will provide much added value or better decisions. In terms of presenting information, probably better to summarize in simpler terms: "On metric one, college X is above/at/below average for it's sector." But then don't just stop there -- we also need to global comparison because people don't get how the sectors vary.
Note that costs are in fact a distribution and presenting average after grants still leaves family very much in the dark if they've no way to know where they're likely to fall on the distribution.
Graduation rates does not suffer from this problem.
Percent of loan repayment is too crude to be useful. It's useful for a banker who may want to finance loans for a student at a given school, but very unclear how this number helps student/family shopping for a college.
Average loan amount is useful.
As important as earning potential is, it's a really stupid number here. Just do a tiny bit of due diligence and you'll see screamingly wide variations across majors, careers, and even within majors. Lawyers, for example, have a certain average starting salary, to be sure, but really big range of variation. Frankly I think putting a single number or even a distribution of incomes next to the name of a school would be nothing but phony quantification. Either that or have a really big footnote explaining statistical significance of differences in means.
What other information would be helpful:
Rather than average loan amount and discount rate what would be useful would be ratios. Tell me (1) list price cost of attendance is X; (2) distribution of discounts is ... and (3) range of debt at graduation is ...
Interesting that you don't really have any room for general comments on doing this at all. You are going to end up diverting an incredible amount of resources toward a project that will in all likelihood produce at best some only moderately useful numbers with huge error bars on them. You will feed into the illusion that choice produces improvement (can you cite any actual evidence?). And you will do absolutely nothing that actually lowers or controls costs, increases graduation rates or lowers indebtedness. In short, not a drop of innovation here. Lots of window dressing, but very little that deserves the name policy.
I'm left wondering why this administration is so confident that "better than the alternative" will continue to be a reason people like me support you.
Does the scorecard cause you to think about things you might not have otherwise considered when choosing a college:
Not in the slightest. It makes me think that whoever made it up has never actually been through the process. It reads more like it is informed by a need to respond to conservative activists who are trying to make hay about higher education. As an Obama supporter and contributor I have to admit it's really a little bit embarrassing to read this as part of the administration's policy proposal. If you can't do better than this, I wonder how bad it would really be to have a republican in the WH as well as in control of congress.
How should this version be modified for 2-year colleges:
Look, it's pretty obvious that there are two issues with two year colleges: (1) to what degree does it lead to successful and timely completion of a four year degree, OR (2) to what degree does it yield serious, usable job training.
So, a start would be to provide rate of students who seek admission to four year who actually graduate from a four year. But really easy to get garbage data on this if you don't set up the categories and the tracking really smart.
On the job side, again, gonna be really serious data quality problems that will likely as not make the information worthless (mostly because you are going to see massive variation from program to program WITHIN schools). That said, let's start with simple "how many people are working in a full-time non-temporary job in or related to the field of their AA degree within X years?"
How should comparison groups for colleges be made? What are important things to consider in grouping institutions together that serve similar students:
Catch-22 here. You are asking people to choose -- if you separate it out too well, the really important thing gets lost: we want people to better understand what the different "rungs" represent. One of the big crimes in higher education is that crappy institutions with minimal value added get to promise people a college degree. And if you only compare within groups each one gets to, in a sense, set the standards. What you need is a tool that more clearly lets people see the payoff differences between the tiers (to the degree there are some).
A most important thing that you'll probably leave out is the effect of what you bring to college on the college outcomes. Huge naivete in college assessment world that the college output has only to do with what the college did. Gigantic effects of origins still at work in higher education. Just be sure your new tool doesn't simply do more to perpetuate the myth.
What search and comparison features would you like the online tool to have:
Something that shows schools in context and behind that groups in context (where does this school sit within its group and where does its group sit in the larger picture).
What should we call this tool? Would a different name better explain the service being provided:
One name would be "Republican Higher Education Policy as Adopted by Obama Administration."
A level headed, but critical, discussion of assessment in higher education by people who "deliver" higher education -- professors.
Sunday, February 5, 2012
Sunday, September 25, 2011
Rubrics, Disenchantment, and Analysis I
There is a tendency, in certain precincts in, and around, higher education, to fethishize rubrics. One gets the impression at conferences and from consultants that arranging something in rows and columns with a few numbers around the edges will call forth the spirit of rational measurement, science even, to descend upon the task at hand. That said, one can acknowledge the heuristic value of rubrics without succumbing to a belief in their magic. Indeed, the critical examination of almost any of the higher education rubrics in current circulation will quickly disenchant, but one need not abandon all hope: if assessment is "here to stay," as some say, it need not be the intellectual train wreck its regional and national champions sometimes seem inclined to produce.
Consider this single item from a rubric used to assess a general education goal in gender:
As is typical of rubric cell content, each of these is "multi-barrelled" -- that is, the description in each cell is asking more than one question at a time. It's not unlike a survey in which respondents are asked, "Are you conservative and in favor of ending welfare?" It's a methodological no-no, and, in general, it defeats the very idea of dis-aggregation (i.e., "what makes up an A?") that a rubric is meant to provide.
In addition, rubrics when they are presented like this are notoriously hard to read. That's not just an aesthetic issue -- failure to communicate effectively leads to misuse of the rubrik (measurement error) and reduces the likelihood of effective constructive critique.
Here is the same information presented in a manner that's more methodologically sound and more intellectually legible:
At the risk of getting ahead of ourselves, there IS a serious problem when these rank ordered categories are used as scores that can be added up and averaged, but we'll save that for another discussion. Too, there is the issue of operationalization -- what does "deep" mean, after all, and how do you distinguish it from not so deep? But this too is for another day.
Let's, for the sake of argument, assume that each of these judgments can be made reliably by competent judges. All told, 4 separate judgments are to be made and each has 3 values. If these knowledges and skills are, in fact, independent (if not, a whole different can of worms), then there are 3 x 3 x 3 x 3 = 81 combinations of ratings possible. Each of these 81 possible assessments is eventually mapped on to1 of 4 ratings. Four combinations are specified, but the other 77 possibilities are not:
Now let us make an (probably invalid) assumption: that each of THESE scores is worth 1, 2 or 3 "points" and then let's calculate the distance between each of the four scores. We use standard Euclidean distance – r=sqrt(x2 + y2) with the categories being: Mastery = 3 3 3 3, Practiced = 2 2 2 3, Introduced = 2 2 2 2, Benchmark = 1 1 1 1
So, how do these categories spread out along the dimension we are measuring here? Mastery, Introduced, and Benchmark are nicely spaced, 2 units apart (and M to B at 4 units). But then we try to fit P in. It's 1.7 units from Mastery and 2.2 from Benchmark, but it's also 1 unit from Introduced. To represent these distances we have to locate it off to the side.
This little exercise suggests that this line of the rubrik is measuring two dimensions.
This should provoke us into thinking about what dimensions of learning are being mixed together in this measurement operation.
It is conventional in this sort of exercise to try to characterize the dimensions in which the items are spread out. Looking back at how we defined the categories we speculate that one dimension might have to do with skill (analysis) and the other knowledge. But Mastery and Practiced were on the same level on analysis. What do we do?
It turns out that the orientation of a diagram like this is arbitrary -- all it is showing us is relative distance. And so we can rotate it like this to show how our assessment categories for this goal relate to one another.
Now you may ask what was the point of this exercise? First, if the point of assessment is to get teachers to think about teaching and learning, and to do so in a manner that applies the same sort of critical thinking skills that we think are important for students to acquire then a careful critique of our assessment methods is absolutely necessary.
Second, this little bit of quick and dirty analysis of a single rubric might actually help people design better rubrics AND to assess the quality of existing rubrics (there's lots more to worry about on these issues, but that's for another time). Maybe, for example, we might conceptualize "introduce" to include knowledge but not skill or vice versa? Maybe we'd think about whether the skill (analysis) is something that should cross GE categories and be expressed in common language. And so on.
Third, this is a first step toward showing why it makes very little sense to take the scores produced by using rubrics like this and then adding them up and averaging them out in order to assess learning. That will be the focus of a subsequent post.
Consider this single item from a rubric used to assess a general education goal in gender:
As is typical of rubric cell content, each of these is "multi-barrelled" -- that is, the description in each cell is asking more than one question at a time. It's not unlike a survey in which respondents are asked, "Are you conservative and in favor of ending welfare?" It's a methodological no-no, and, in general, it defeats the very idea of dis-aggregation (i.e., "what makes up an A?") that a rubric is meant to provide.
In addition, rubrics when they are presented like this are notoriously hard to read. That's not just an aesthetic issue -- failure to communicate effectively leads to misuse of the rubrik (measurement error) and reduces the likelihood of effective constructive critique.
Here is the same information presented in a manner that's more methodologically sound and more intellectually legible:
At the risk of getting ahead of ourselves, there IS a serious problem when these rank ordered categories are used as scores that can be added up and averaged, but we'll save that for another discussion. Too, there is the issue of operationalization -- what does "deep" mean, after all, and how do you distinguish it from not so deep? But this too is for another day.
Let's, for the sake of argument, assume that each of these judgments can be made reliably by competent judges. All told, 4 separate judgments are to be made and each has 3 values. If these knowledges and skills are, in fact, independent (if not, a whole different can of worms), then there are 3 x 3 x 3 x 3 = 81 combinations of ratings possible. Each of these 81 possible assessments is eventually mapped on to1 of 4 ratings. Four combinations are specified, but the other 77 possibilities are not:
Now let us make an (probably invalid) assumption: that each of THESE scores is worth 1, 2 or 3 "points" and then let's calculate the distance between each of the four scores. We use standard Euclidean distance – r=sqrt(x2 + y2) with the categories being: Mastery = 3 3 3 3, Practiced = 2 2 2 3, Introduced = 2 2 2 2, Benchmark = 1 1 1 1
So, how do these categories spread out along the dimension we are measuring here? Mastery, Introduced, and Benchmark are nicely spaced, 2 units apart (and M to B at 4 units). But then we try to fit P in. It's 1.7 units from Mastery and 2.2 from Benchmark, but it's also 1 unit from Introduced. To represent these distances we have to locate it off to the side.
This little exercise suggests that this line of the rubrik is measuring two dimensions.
This should provoke us into thinking about what dimensions of learning are being mixed together in this measurement operation.
It is conventional in this sort of exercise to try to characterize the dimensions in which the items are spread out. Looking back at how we defined the categories we speculate that one dimension might have to do with skill (analysis) and the other knowledge. But Mastery and Practiced were on the same level on analysis. What do we do?
It turns out that the orientation of a diagram like this is arbitrary -- all it is showing us is relative distance. And so we can rotate it like this to show how our assessment categories for this goal relate to one another.
Now you may ask what was the point of this exercise? First, if the point of assessment is to get teachers to think about teaching and learning, and to do so in a manner that applies the same sort of critical thinking skills that we think are important for students to acquire then a careful critique of our assessment methods is absolutely necessary.
Second, this little bit of quick and dirty analysis of a single rubric might actually help people design better rubrics AND to assess the quality of existing rubrics (there's lots more to worry about on these issues, but that's for another time). Maybe, for example, we might conceptualize "introduce" to include knowledge but not skill or vice versa? Maybe we'd think about whether the skill (analysis) is something that should cross GE categories and be expressed in common language. And so on.
Third, this is a first step toward showing why it makes very little sense to take the scores produced by using rubrics like this and then adding them up and averaging them out in order to assess learning. That will be the focus of a subsequent post.
Labels:
assessment,
assessment rubrics,
measurement,
methodology
Sunday, August 14, 2011
What Will "Assessment 2.0" Look Like? A Proposal
The most serious flaw in assessment as now practiced is the premise that it is something that teachers are not interested in, do not want to do, have not been doing, etc. A word that comes up a lot in connection with assessment is "accountability," but most folks who use the word don't take the time to be explicit about just who is supposed to be accountable to whom for what. When someone does get beyond just parroting the word, the most common interpretation seems to be "we need to hold teachers accountable."
We have some news for those who have discovered assessment. Teachers -- lecturers, instructors, professors -- have long been interested in what works and what doesn't in the classroom. Those who would appoint themselves guardians of learning have a nasty habit of trotting out stereotypes of the worst professor ever and, in a classic example of question begging, concluding that such figures dominate the academy and represent a threat to the future of higher education.
But rather than argue about that, here's a proposal for what the next stage in assessment might look like.
Given that most professors and most departments are actually interested in student learning and in how to maximize it -- this is, after all, the vocation these folks have chosen -- the resources that have been pumped into assessment projects should be put at the service of the faculty. Throughout Assessment 1.0 the dominant pattern is for an office of assessment to be in the driver's seat, more or less dictating to faculty (generally relaying what had been dictated to them by accreditation agencies) how and when assessment would happen. Many faculty found the methods wanting and the tasks tedious and pointless, but most went along -- at some institutions more willingly and at some less. The interaction between faculty and assessment offices generally came down to the latter making work for the former without the former seeing much in the way of benefits.
That's unfortunate because there are lots of potential benefits for us as instructors. But to realize them, we need to turn the tables. The basic premise of Assessment 2.0 should be (1) that it be faculty driven and (2) that assessment offices work for the faculty, rather than the other way round. Assessment offices should think of themselves as a support service for the academic program rather than a support service for a regulatory body that oversees the academic program from the outside. The main job of assessment offices should be to make a part of the job that faculty do, as professionals practicing their craft, easier. A part of what professionals do is self monitor and mutually monitor outcomes. As faculty, we need to think about what information will help us to make micro-, meso-, and macro-adjustments in our practice that will improve the outcomes we are collectively trying to achieve.
And the services of our assessment offices should be available to us to obtain it. We need to put the focus back on this side of the operation and shift away from the idea that the primary motivation behind assessment is to prove something to outsiders. Even the rhetoric from the accreditation agencies, if you slow the tape down and listen, resonates with this: they demand evidence that assessment is happening, that program adjustments happen in response to it, and so on. Where they are wrong is in their ignorant insistence that such things were not already happening.
The assessment industry did not invent assessment -- they simply codified it and figured out how to make a living off of doing it instead of being involved directly in educating.
We have some news for those who have discovered assessment. Teachers -- lecturers, instructors, professors -- have long been interested in what works and what doesn't in the classroom. Those who would appoint themselves guardians of learning have a nasty habit of trotting out stereotypes of the worst professor ever and, in a classic example of question begging, concluding that such figures dominate the academy and represent a threat to the future of higher education.
But rather than argue about that, here's a proposal for what the next stage in assessment might look like.
Given that most professors and most departments are actually interested in student learning and in how to maximize it -- this is, after all, the vocation these folks have chosen -- the resources that have been pumped into assessment projects should be put at the service of the faculty. Throughout Assessment 1.0 the dominant pattern is for an office of assessment to be in the driver's seat, more or less dictating to faculty (generally relaying what had been dictated to them by accreditation agencies) how and when assessment would happen. Many faculty found the methods wanting and the tasks tedious and pointless, but most went along -- at some institutions more willingly and at some less. The interaction between faculty and assessment offices generally came down to the latter making work for the former without the former seeing much in the way of benefits.
That's unfortunate because there are lots of potential benefits for us as instructors. But to realize them, we need to turn the tables. The basic premise of Assessment 2.0 should be (1) that it be faculty driven and (2) that assessment offices work for the faculty, rather than the other way round. Assessment offices should think of themselves as a support service for the academic program rather than a support service for a regulatory body that oversees the academic program from the outside. The main job of assessment offices should be to make a part of the job that faculty do, as professionals practicing their craft, easier. A part of what professionals do is self monitor and mutually monitor outcomes. As faculty, we need to think about what information will help us to make micro-, meso-, and macro-adjustments in our practice that will improve the outcomes we are collectively trying to achieve.
And the services of our assessment offices should be available to us to obtain it. We need to put the focus back on this side of the operation and shift away from the idea that the primary motivation behind assessment is to prove something to outsiders. Even the rhetoric from the accreditation agencies, if you slow the tape down and listen, resonates with this: they demand evidence that assessment is happening, that program adjustments happen in response to it, and so on. Where they are wrong is in their ignorant insistence that such things were not already happening.
The assessment industry did not invent assessment -- they simply codified it and figured out how to make a living off of doing it instead of being involved directly in educating.
Thursday, August 11, 2011
Too Bad Higher Education "Experts" and Vendors aren't Graded
I was inspired by a TeachSoc post from Kathe Lowney today to have a look at two articles in the Chronicle of Higher Education on computer essay grading.
The articles are "Professors Cede Grading Power to Outsiders—Even Computers" and "Can Software Make the Grade?"
My Review: A typical Chronicle hack job to my mind. Articles like this remind me of National Enquirer. Author makes little attempt to critically assess comments from his sources and gives little weight to contrary information (failing to infer, for example, anything from reported fact that in six years of marketing, almost no one has bought into the computer grading product mentioned). He jumps on grade inflation bandwagon instead of offering an analytic take on it. In typical COHE fashion he sets up false dichotomies and debates between advocates and defenders as if there is a big divide down the middle of higher education. In effect, articles like this are just product placement -- hopefully without kickbacks -- and "if someone says it then it's a usable quote" journalism. As with many COHE articles, it reflects journalism that's more in touch with the higher education industry than with higher education. It's mediocre work such as this that makes me let my subscription lapse every year or so. It's interesting how COHE seems to have no qualms at all about trashing educators and educational institutions but only ever so rarely do they seem to take an even gentle critical look at education vendors.
My Review: A typical Chronicle hack job to my mind. Articles like this remind me of National Enquirer. Author makes little attempt to critically assess comments from his sources and gives little weight to contrary information (failing to infer, for example, anything from reported fact that in six years of marketing, almost no one has bought into the computer grading product mentioned). He jumps on grade inflation bandwagon instead of offering an analytic take on it. In typical COHE fashion he sets up false dichotomies and debates between advocates and defenders as if there is a big divide down the middle of higher education. In effect, articles like this are just product placement -- hopefully without kickbacks -- and "if someone says it then it's a usable quote" journalism. As with many COHE articles, it reflects journalism that's more in touch with the higher education industry than with higher education. It's mediocre work such as this that makes me let my subscription lapse every year or so. It's interesting how COHE seems to have no qualms at all about trashing educators and educational institutions but only ever so rarely do they seem to take an even gentle critical look at education vendors.
On the accompanying "compare yourself to the computer" article : I think I'd fire a TA who graded like that -- the words "capitalism" and "rationality" showing up constitute "concepts related to him" and an answer on Marx where "expelled for advocating revolution" = "significance for social science"? I scored them 4 and 2 and that was generous. I'd be mighty disappointed if I were the makers of that software and this is how my product placement in COHE turned out -- would anyone buy it based on this portrayal?!
Tuesday, March 22, 2011
The Rubrikization of Higher Education
The rubricization of education has always rubbed me the wrong way but I’ve never been able to put my finger on concrete flaws beyond the obvious. This past January I attended the AAC&U conference in San Francisco. A few more problems became clear.
There are three obvious methodological/measurement problems that have long stood out:
1. Almost every rubric I have ever seen has exhibited gads of multi-dimensionality in the different skills/items/categories/rows. Another way to say this is that the rows typically posed double or multi-barreled questions to the evaluator. Or, even if the construct named in the row was simple, the description of the different scale levels would be multi-dimensional. Example:
| Category | Advanced (4) | Competent (3) | Developing (2) | Underdeveloped (1) |
| Structure | Sections fit together in logical sequence; claims, evidence, analysis, conclusions distinguished; logic of argument telescoped and reviewed |
One argument that this is not a problem is that all the things listed here typically go together and that they are all indicators of the same underlying skill. Maybe. But it seems to be a stretch that all these skills nicely fall into a simple four level linear scale.
2. The second problem here is just that four point scale. What evidence is there to support the idea that “Advanced” level structure is two times as much structure (or as much skill) as “Developing”? This does not matter much when we are simply looking at these four levels, but the first thing that that folks with just a little quantitative skill do is come up with average ratings for a group of students on a skill rating like this.
Let us be clear: computing the average of a scale that has not been shown to have the arithmetic properties of what we call an interval scale PRODUCES MEANINGLESS RESULTS.
3. The third problem with rubriks like this is that the items (rows) are not necessarily exhaustive or mutually exclusive. In other words, they do not always include all the components of learning that might be (or should be) happening and the individual items often tap into the same underlying skill. The former is a substantive problem to be solved by better conversations about the goals of education. The latter, though, lead to bad data. Suppose three items X, Y, and Z are listed in a rubric and that the elaborate operationalizations of the different levels of these involve underlying skills a, b, c, d, and e.
| Category | Advanced (4) | Competent (3) | Developing (2) | Underdeveloped (1) |
| X | Blah blah blah {a} blah blah blah {c} | |||
| Y | Blah blah blah {b} blah blah blah {c} | |||
| Z | Blah blah blah {d} blah blah blah {a} blah blah blah {e} blah blah blah {c} |
Where we’ve put in curly brackets the underlying skill that the description “blah blah blah” refers to. In this rubrik, skills a and b get counted twice, skill c three times. When data is aggregated, success on a, b, or c will easily mask lack of progress on d or e.
4. But here is the most serious problem of rubricization. It completely drives out of the teaching and learning process any response to individual variations in understanding. The role of the teacher as offering constructive criticism about the wide range of variability in learning is driven out in favor of a set of categories.
One great irony in this is that so many of the champions of this approach to educational reform are the very folks who preach about variability of learning styles.
Another is the high level of concern about students who “fall between the cracks.” Here we are developing a system with explicitly designed cracks between which they can fall.
Yet another is that a mantra of the rubrik crowd is “evidence based” and “data driven” decisions. And yet the very devices that lie at the heart of the enterprise are custom-built to degrade information and result in misleading data.
The fundamental absence of critical thinking in the rubrik/assessment literature – and total lack of interest in critical discourse about these techniques – is the final irony.
One can conclude that what we have here is a bunch of middle-brow thinkers designing a system that will maximize the production of people like themselves and guarantee their own employment in higher education industry. If only there were some evidence that this is what the world will need in the 21st century.
Saturday, December 4, 2010
Coming Soon to a Classroom Near You?
Some rambling thoughts on a fascinating set of articles about measuring teaching.
Today's NYT carried two stories -- on on page 1 -- about new techniques being used to evaluate K-12 teachers. The news in the stories concerns two things: existence of a very large program for measuring educational effectiveness in schools and the central role of video-taping teachers teaching in that program.
Local readers' radar might ponder the resonance between programs like this and higher education assessment and higher education "learning and teaching centers" and the individuals and organizations who live off, rather than for, education.
The first story ("Teacher Ratings Get New Look, Pushed by a Rich Watcher") highlights Bill Gates' (via the Gates Foundation) interest in a gigantic project measuring the "value added" by teachers through multi-mode assessment. Among other tools : videos of instruction that are scored by experts.
"Interesting" is the fact that one of the movers and shakers in the project is none other than Educational Testing Services. And so this represents yet another opportunity for that organization to live off, rather than for, education in the U.S. Other contractors are mentioned in the story too -- as has been true of the assessment movement more generally, a big part of the driving force seems to be entrepreneurs who, after persuading you that you need to do something are more than happy to sell you the equipment needed to collect the data and then expertise to evaluate it.
The second article, "Video Eye Aimed at Teachers in 7 School Systems," describes some 3,000 teachers who are a part of the first phase of this search for new methods to evaluate teachers. Each will have several hours of teaching video-taped and the tapes will be assessed by experts using a number of carefully validated protocols.
The first article, describing the scope of the project, notes that the rating of 24,000 video-taped lessons will come to something like 64,000 hours of video watching. On a full-time basis that represents 32 person years of work. At 180 days/year, that's about 44 years of teaching. The article suggests the costs to a school district will be about $1.5 million up front and then $800,000 per year.
I wonder if anyone has assessed the value of the information produced.
In the middle of the report there is a line about how this is a step forward because rather than having the principal observe once or twice during the year, outside experts (using scientific protocols) can observe up to a half dozen times. This suggests an interesting phenomenon: in the name of standardization and objectivity, we deskill and depersonalize (among other things).
In one paper on value added modeling (VAM), by an ETS staff person (Braun 2004, 17), one finds this argument: (1) quantitative evaluation of teaching is here to stay; (2) evaluation of gains is preferable to just measuring year-end performance; (3) we have to think what would get used if not this; (4) therefore, use VAM even if it has real limitations. Another, by a Michigan State University economist concludes (about VAM):
Resources
Amrein-Beardsley, Audrey. 2008. "Methodological Concerns About the Education Value-Added Assessment System." Educational Researcher, Vol. 37, No. 2, pp. 65–75
Braun, Henry. 2004. "VALUE-ADDED MODELING: WHAT DOES DUE DILIGENCE REQUIRE?"
Rand Corporation. 2007. "The Promise and Peril of Using Value-Added Modeling to Measure Teacher Effectiveness"
Reckase, Mark D. 2004. "Measurement Issues Associated with Value-added Methods"
Wikipedia. "Value Added Modeling"
Today's NYT carried two stories -- on on page 1 -- about new techniques being used to evaluate K-12 teachers. The news in the stories concerns two things: existence of a very large program for measuring educational effectiveness in schools and the central role of video-taping teachers teaching in that program.
Local readers' radar might ponder the resonance between programs like this and higher education assessment and higher education "learning and teaching centers" and the individuals and organizations who live off, rather than for, education.
The first story ("Teacher Ratings Get New Look, Pushed by a Rich Watcher") highlights Bill Gates' (via the Gates Foundation) interest in a gigantic project measuring the "value added" by teachers through multi-mode assessment. Among other tools : videos of instruction that are scored by experts.
"Interesting" is the fact that one of the movers and shakers in the project is none other than Educational Testing Services. And so this represents yet another opportunity for that organization to live off, rather than for, education in the U.S. Other contractors are mentioned in the story too -- as has been true of the assessment movement more generally, a big part of the driving force seems to be entrepreneurs who, after persuading you that you need to do something are more than happy to sell you the equipment needed to collect the data and then expertise to evaluate it.
The second article, "Video Eye Aimed at Teachers in 7 School Systems," describes some 3,000 teachers who are a part of the first phase of this search for new methods to evaluate teachers. Each will have several hours of teaching video-taped and the tapes will be assessed by experts using a number of carefully validated protocols.
The first article, describing the scope of the project, notes that the rating of 24,000 video-taped lessons will come to something like 64,000 hours of video watching. On a full-time basis that represents 32 person years of work. At 180 days/year, that's about 44 years of teaching. The article suggests the costs to a school district will be about $1.5 million up front and then $800,000 per year.
I wonder if anyone has assessed the value of the information produced.
In the middle of the report there is a line about how this is a step forward because rather than having the principal observe once or twice during the year, outside experts (using scientific protocols) can observe up to a half dozen times. This suggests an interesting phenomenon: in the name of standardization and objectivity, we deskill and depersonalize (among other things).
In one paper on value added modeling (VAM), by an ETS staff person (Braun 2004, 17), one finds this argument: (1) quantitative evaluation of teaching is here to stay; (2) evaluation of gains is preferable to just measuring year-end performance; (3) we have to think what would get used if not this; (4) therefore, use VAM even if it has real limitations. Another, by a Michigan State University economist concludes (about VAM):
We are looking at the educational system through a poor quality lens. The real world is probably more orderly than it appears from the analyses of noisy data (Reckase 2004, 7).
Resources
Amrein-Beardsley, Audrey. 2008. "Methodological Concerns About the Education Value-Added Assessment System." Educational Researcher, Vol. 37, No. 2, pp. 65–75
Braun, Henry. 2004. "VALUE-ADDED MODELING: WHAT DOES DUE DILIGENCE REQUIRE?"
Rand Corporation. 2007. "The Promise and Peril of Using Value-Added Modeling to Measure Teacher Effectiveness"
Reckase, Mark D. 2004. "Measurement Issues Associated with Value-added Methods"
Wikipedia. "Value Added Modeling"
Monday, September 6, 2010
Closing the Loop in Practice: Does Assessment Get Assessment?
At a liberal arts college with which I am familiar, the administration recently distributed "syllabus guidelines" with 34 items for inclusion on course syllabi. Faculty leaders balked and asked for clarification: which of the 34 items were mandates (and from whom on what authority) and which were someone's "good idea"? The response was that guidelines are merely guidelines and most of the content were indeed good ideas. Most were.
A subsequent examination of a sample of syllabi revealed that most syllabi did not contain all 34. More specifically, there was not universal inclusion of several that, apparently, are important for accreditation purposes.
The semester has begun. The syllabi are printed. The administration disseminated the guidelines -- their obligation is fulfilled. If faculty choose not to comply, that's their decision. Overall, the situation is alarming because the school could appear to be non-compliant to its accreditors. And it's the faculty's fault. And folks are wondering how to fix it.
THIS COULD BE TURNED INTO SOMETHING POSITIVE, a shining example of assessment, closing the loop, and evidence-based change.
But first, WAIT A MINUTE! Do faculty get to say "We told them what to do; if they can't comply and don't learn, it's not my fault."? Of course not. If students aren't learning, faculty are doing something wrong. Lack of learning = feedback, and feedback must lead to change.
Here we have a case of an institution ignoring unambiguous feedback. The feedback is simple: the promulgation of a list of 34 things one should do on a syllabus does not produce the uniform inclusion of the small handful of actually really important things to include on a syllabus. That's it; that's what the evidence tells you. It doesn't tell you faculty are bad; it tells you that this method of changing what syllabi look like was ineffective.
Never mind that any good teacher knows that you cannot motivate change with a list of 34 fixes.
The correct response? Close the loop: listen, learn, change the way syllabus guidelines are handled.
The unfortunate thing here is that folks who know (faculty) brought this immediately to the attention of the folks in charge. Faculty noted that the list was too long, its provenance ambiguous, its authority unclear, its applicability variable, its tone insulting. A solution was suggested. All this was met with, basically, a brush off -- they're just guidelines not requirements, what's the big deal?
And, it turns out, that is precisely how faculty understood them. No need for alarm. Some adjusted their syllabi to some of the suggestions in the guidelines. But apparently, the faculty didn't all implement a few of the guidelines that really do matter (to someone). Arrrrrrrgh.
And now for a little forward looking fantasy of what the outcome of this situation COULD be.
Educators really committed to the stated goals of assessment would see in this affair an opportunity for an achievement they could boast about. Those committed to one directional, top-down, assessor-centered, non-interactive, deaf-to-feedback approaches will see in it only faculty reluctance to get with the program.
One lesson learned here is that institutional processes need adjustment. The amount of faculty and administrative time, emotional energy, and the augmentation of frustration and mistrust that this little thing has engendered was a phenomenal waste of precious institutional resources. Alas, accountability for THIS is unlikely ever to be reckoned.
A subsequent examination of a sample of syllabi revealed that most syllabi did not contain all 34. More specifically, there was not universal inclusion of several that, apparently, are important for accreditation purposes.
The semester has begun. The syllabi are printed. The administration disseminated the guidelines -- their obligation is fulfilled. If faculty choose not to comply, that's their decision. Overall, the situation is alarming because the school could appear to be non-compliant to its accreditors. And it's the faculty's fault. And folks are wondering how to fix it.
THIS COULD BE TURNED INTO SOMETHING POSITIVE, a shining example of assessment, closing the loop, and evidence-based change.
But first, WAIT A MINUTE! Do faculty get to say "We told them what to do; if they can't comply and don't learn, it's not my fault."? Of course not. If students aren't learning, faculty are doing something wrong. Lack of learning = feedback, and feedback must lead to change.
Here we have a case of an institution ignoring unambiguous feedback. The feedback is simple: the promulgation of a list of 34 things one should do on a syllabus does not produce the uniform inclusion of the small handful of actually really important things to include on a syllabus. That's it; that's what the evidence tells you. It doesn't tell you faculty are bad; it tells you that this method of changing what syllabi look like was ineffective.
Never mind that any good teacher knows that you cannot motivate change with a list of 34 fixes.
The correct response? Close the loop: listen, learn, change the way syllabus guidelines are handled.
The unfortunate thing here is that folks who know (faculty) brought this immediately to the attention of the folks in charge. Faculty noted that the list was too long, its provenance ambiguous, its authority unclear, its applicability variable, its tone insulting. A solution was suggested. All this was met with, basically, a brush off -- they're just guidelines not requirements, what's the big deal?
And, it turns out, that is precisely how faculty understood them. No need for alarm. Some adjusted their syllabi to some of the suggestions in the guidelines. But apparently, the faculty didn't all implement a few of the guidelines that really do matter (to someone). Arrrrrrrgh.
And now for a little forward looking fantasy of what the outcome of this situation COULD be.
Since administrations and the assessment industry are apparently NOT really ready to adopt the underlying premise of assessment -- pay attention to feedback and change accordingly -- the faculty will.
From now on, only the faculty will disseminate syllabi guidelines. They will very clearly distinguish between legally mandated content, accreditation relevant functionality, college-specific custom and standards, and good pedagogical practice in general. They will invite all parties who become aware of syllabi-related mandates (or new good ideas) to communicate them to the faculty's educational policy committee for consideration for inclusion in their next semester's guidelines.
Those guidelines will explicitly articulate general goals (exactly which ones to be determined) such as syllabi are to be interesting documents that are useful to students and that permit colleagues to get a sense of what a course is about and at what level it is being taught as well as suggestions of particular features, boilerplate and examples that might be useful, and fully explained required items. They will include examples of an array of syllabi that explicitly demonstrate a variety of forms that meet their standards. And, all suggestions will be referenced where possible and requirements will be documented in terms of on what authority they are an obligation.
For assessment purposes the faculty will adapt* any externally supplied "rubrics" to their own intellectually and pedagogically defensible standards and practices and encourage our colleagues to make use of these college-specific tools in developing their syllabi.
From now on, only the faculty will disseminate syllabi guidelines. They will very clearly distinguish between legally mandated content, accreditation relevant functionality, college-specific custom and standards, and good pedagogical practice in general. They will invite all parties who become aware of syllabi-related mandates (or new good ideas) to communicate them to the faculty's educational policy committee for consideration for inclusion in their next semester's guidelines.
Those guidelines will explicitly articulate general goals (exactly which ones to be determined) such as syllabi are to be interesting documents that are useful to students and that permit colleagues to get a sense of what a course is about and at what level it is being taught as well as suggestions of particular features, boilerplate and examples that might be useful, and fully explained required items. They will include examples of an array of syllabi that explicitly demonstrate a variety of forms that meet their standards. And, all suggestions will be referenced where possible and requirements will be documented in terms of on what authority they are an obligation.
For assessment purposes the faculty will adapt* any externally supplied "rubrics" to their own intellectually and pedagogically defensible standards and practices and encourage our colleagues to make use of these college-specific tools in developing their syllabi.
Educators really committed to the stated goals of assessment would see in this affair an opportunity for an achievement they could boast about. Those committed to one directional, top-down, assessor-centered, non-interactive, deaf-to-feedback approaches will see in it only faculty reluctance to get with the program.
One lesson learned here is that institutional processes need adjustment. The amount of faculty and administrative time, emotional energy, and the augmentation of frustration and mistrust that this little thing has engendered was a phenomenal waste of precious institutional resources. Alas, accountability for THIS is unlikely ever to be reckoned.
Subscribe to:
Posts (Atom)





