Sunday, August 14, 2011

What Will "Assessment 2.0" Look Like? A Proposal

The most serious flaw in assessment as now practiced is the premise that it is something that teachers are not interested in, do not want to do, have not been doing, etc.  A word that comes up a lot in connection with assessment is "accountability," but most folks who use the word don't take the time to be explicit about just who is supposed to be accountable to whom for what.  When someone does get beyond just parroting the word, the most common interpretation seems to be "we need to hold teachers accountable."

We have some news for those who have discovered assessment.  Teachers -- lecturers, instructors, professors -- have long been interested in what works and what doesn't in the classroom.  Those who would appoint themselves guardians of learning have a nasty habit of trotting out stereotypes of the worst professor ever and, in a classic example of question begging, concluding that such figures dominate the academy and represent a threat to the future of higher education.

But rather than argue about that, here's a proposal for what the next stage in assessment might look like.

Given that most professors and most departments are actually interested in student learning and in how to maximize it -- this is, after all, the vocation these folks have chosen -- the resources that have been pumped into assessment projects should be put at the service of the faculty.  Throughout Assessment 1.0 the dominant pattern is for an office of assessment to be in the driver's seat, more or less dictating to faculty (generally relaying what had been dictated to them by accreditation agencies) how and when assessment would happen.  Many faculty found the methods wanting and the tasks tedious and pointless, but most went along -- at some institutions more willingly and at some less.  The interaction between faculty and assessment offices generally came down to the latter making work for the former without the former seeing much in the way of benefits.

That's unfortunate because there are lots of potential benefits for us as instructors.  But to realize them, we need to turn the tables.  The basic premise of Assessment 2.0 should be (1) that it be faculty driven and (2) that assessment offices work for the faculty, rather than the other way round.  Assessment offices should think of themselves as a support service for the academic program rather than a support service for a regulatory body that oversees the academic program from the outside.  The main job of assessment offices should be to make a part of the job that faculty do, as professionals practicing their craft, easier.  A part of what professionals do is self monitor and mutually monitor outcomes.  As faculty, we need to think about what information will help us to make micro-, meso-, and macro-adjustments in our practice that will improve the outcomes we are collectively trying to achieve.

And the services of our assessment offices should be available to us to obtain it.  We need to put the focus back on this side of the operation and shift away from the idea that the primary motivation behind assessment is to prove something to outsiders.  Even the rhetoric from the accreditation agencies, if you slow the tape down and listen, resonates with this: they demand evidence that assessment is happening, that program adjustments happen in response to it, and so on.  Where they are wrong is in their ignorant insistence that such things were not already happening.

The assessment industry did not invent assessment -- they simply codified it and figured out how to make a living off of doing it instead of being involved directly in educating.

Thursday, August 11, 2011

Too Bad Higher Education "Experts" and Vendors aren't Graded


I was inspired by a TeachSoc post from Kathe Lowney today to have a look at two articles in the Chronicle of Higher Education on computer essay grading.

The articles are "Professors Cede Grading Power to Outsiders—Even Computers" and "Can Software Make the Grade?"

My Review: A typical Chronicle hack job to my mind.  Articles like this remind me of National Enquirer.  Author makes little attempt to critically assess comments from his sources and gives little weight  to contrary information (failing to infer, for example, anything from reported fact that in six years of marketing, almost no one has bought into the computer grading product mentioned).  He jumps on grade inflation bandwagon instead of offering an analytic take on it.  In typical COHE fashion he sets up false dichotomies and debates between advocates and defenders as if there is a big divide down the middle of higher education.  In effect, articles like this are just product placement -- hopefully without kickbacks -- and "if someone says it then it's a usable quote" journalism.  As with many COHE articles, it reflects journalism that's more in touch with the higher education industry than with higher education.   It's mediocre work such as  this that makes me let my subscription lapse every year or so.  It's interesting how COHE  seems to have no qualms at all about trashing educators and educational institutions but only ever so rarely do they seem to take an even gentle critical look at education vendors.

On the accompanying "compare yourself to the computer" article : I think I'd fire a TA who graded like that -- the words "capitalism" and "rationality" showing up constitute "concepts related to him" and an answer on Marx where "expelled for advocating revolution" = "significance for social science"?  I scored them 4 and 2 and that was generous.   I'd be mighty disappointed if I were the makers of that software and this is how my product placement in COHE turned out -- would anyone buy it based on this portrayal?! 





Tuesday, March 22, 2011

The Rubrikization of Higher Education

The rubricization of education has always rubbed me the wrong way but I’ve never been able to put my finger on concrete flaws beyond the obvious. This past January I attended the AAC&U conference in San Francisco. A few more problems became clear.

There are three obvious methodological/measurement problems that have long stood out:

1. Almost every rubric I have ever seen has exhibited gads of multi-dimensionality in the different skills/items/categories/rows. Another way to say this is that the rows typically posed double or multi-barreled questions to the evaluator. Or, even if the construct named in the row was simple, the description of the different scale levels would be multi-dimensional. Example:

Category Advanced (4) Competent (3) Developing (2) Underdeveloped (1)
Structure Sections fit together in logical sequence; claims, evidence, analysis, conclusions distinguished; logic of argument telescoped and reviewed

One argument that this is not a problem is that all the things listed here typically go together and that they are all indicators of the same underlying skill. Maybe. But it seems to be a stretch that all these skills nicely fall into a simple four level linear scale.

2. The second problem here is just that four point scale. What evidence is there to support the idea that “Advanced” level structure is two times as much structure (or as much skill) as “Developing”? This does not matter much when we are simply looking at these four levels, but the first thing that that folks with just a little quantitative skill do is come up with average ratings for a group of students on a skill rating like this.

Let us be clear: computing the average of a scale that has not been shown to have the arithmetic properties of what we call an interval scale PRODUCES MEANINGLESS RESULTS.

3. The third problem with rubriks like this is that the items (rows) are not necessarily exhaustive or mutually exclusive. In other words, they do not always include all the components of learning that might be (or should be) happening and the individual items often tap into the same underlying skill. The former is a substantive problem to be solved by better conversations about the goals of education. The latter, though, lead to bad data. Suppose three items X, Y, and Z are listed in a rubric and that the elaborate operationalizations of the different levels of these involve underlying skills a, b, c, d, and e.

Category Advanced (4) Competent (3) Developing (2) Underdeveloped (1)
X Blah blah blah {a} blah blah blah {c}
Y Blah blah blah {b} blah blah blah {c}
Z Blah blah blah {d} blah blah blah {a} blah blah blah {e} blah blah blah {c}

Where we’ve put in curly brackets the underlying skill that the description “blah blah blah” refers to. In this rubrik, skills a and b get counted twice, skill c three times. When data is aggregated, success on a, b, or c will easily mask lack of progress on d or e.

4. But here is the most serious problem of rubricization. It completely drives out of the teaching and learning process any response to individual variations in understanding. The role of the teacher as offering constructive criticism about the wide range of variability in learning is driven out in favor of a set of categories.

One great irony in this is that so many of the champions of this approach to educational reform are the very folks who preach about variability of learning styles.

Another is the high level of concern about students who “fall between the cracks.” Here we are developing a system with explicitly designed cracks between which they can fall.

Yet another is that a mantra of the rubrik crowd is “evidence based” and “data driven” decisions. And yet the very devices that lie at the heart of the enterprise are custom-built to degrade information and result in misleading data.

The fundamental absence of critical thinking in the rubrik/assessment literature – and total lack of interest in critical discourse about these techniques – is the final irony.

One can conclude that what we have here is a bunch of middle-brow thinkers designing a system that will maximize the production of people like themselves and guarantee their own employment in higher education industry. If only there were some evidence that this is what the world will need in the 21st century.

Saturday, December 4, 2010

Coming Soon to a Classroom Near You?

Some rambling thoughts on a fascinating set of articles about measuring teaching.

Today's NYT carried two stories -- on on page 1 -- about new techniques being used to evaluate K-12 teachers.  The news in the stories concerns two things: existence of a very large program for measuring educational effectiveness in schools and the central role of video-taping teachers teaching in that program.

Local readers' radar might ponder the resonance between programs like this and higher education assessment and higher education "learning and teaching centers" and the individuals and organizations who live off, rather than for, education.

The first story ("Teacher Ratings Get New Look, Pushed by a Rich Watcher") highlights Bill Gates' (via the Gates Foundation) interest in a gigantic project measuring the "value added" by teachers through multi-mode assessment. Among other tools : videos of instruction that are scored by experts.

"Interesting" is the fact that one of the movers and shakers in the project is none other than Educational Testing Services. And so this represents yet another opportunity for that organization to live off, rather than for, education in the U.S. Other contractors are mentioned in the story too -- as has been true of the assessment movement more generally, a big part of the driving force seems to be entrepreneurs who, after persuading you that you need to do something are more than happy to sell you the equipment needed to collect the data and then expertise to evaluate it.

The second article, "Video Eye Aimed at Teachers in 7 School Systems," describes some 3,000 teachers who are a part of the first phase of this search for new methods to evaluate teachers. Each will have several hours of teaching video-taped and the tapes will be assessed by experts using a number of carefully validated protocols.

The first article, describing the scope of the project, notes that the rating of 24,000 video-taped lessons will come to something like 64,000 hours of video watching. On a full-time basis that represents 32 person years of work. At 180 days/year, that's about 44 years of teaching.  The article suggests the costs to a school district will be about $1.5 million up front and then $800,000 per year.

I wonder if anyone has assessed the value of the information produced.

In the middle of the report there is a line about how this is a step forward because rather than having the principal observe once or twice during the year, outside experts (using scientific protocols) can observe up to a half dozen times.  This suggests an interesting phenomenon: in the name of standardization and objectivity, we deskill and depersonalize (among other things).

In one paper on value added modeling (VAM), by an ETS staff person (Braun 2004, 17), one finds this argument: (1) quantitative evaluation of teaching is here to stay; (2) evaluation of gains is preferable to just measuring year-end performance; (3) we have to think what would get used if not this; (4) therefore, use VAM even if it has real limitations. Another, by a Michigan State University economist concludes (about VAM):
We are looking at the educational system through a poor quality lens. The real world is probably more orderly than it appears from the analyses of noisy data (Reckase 2004, 7).

Resources

Amrein-Beardsley, Audrey. 2008. "Methodological Concerns About the Education Value-Added Assessment System." Educational Researcher, Vol. 37, No. 2, pp. 65–75

Braun, Henry. 2004. "VALUE-ADDED MODELING: WHAT DOES DUE DILIGENCE REQUIRE?"

Rand Corporation. 2007. "The Promise and Peril of Using Value-Added Modeling to Measure Teacher Effectiveness"

Reckase, Mark D. 2004. "Measurement Issues Associated with Value-added Methods"

Wikipedia. "Value Added Modeling"

Monday, September 6, 2010

Closing the Loop in Practice: Does Assessment Get Assessment?

At a liberal arts college with which I am familiar, the administration recently distributed "syllabus guidelines" with 34 items for inclusion on course syllabi. Faculty leaders balked and asked for clarification: which of the 34 items were mandates (and from whom on what authority) and which were someone's "good idea"? The response was that guidelines are merely guidelines and most of the content were indeed good ideas. Most were.

A subsequent examination of a sample of syllabi revealed that most syllabi did not contain all 34. More specifically, there was not universal inclusion of several that, apparently, are important for accreditation purposes. 

The semester has begun.  The syllabi are printed.  The administration disseminated the guidelines -- their obligation is fulfilled.  If faculty choose not to comply, that's their decision. Overall, the situation is alarming because the school could appear to be non-compliant to its accreditors.  And it's the faculty's fault.  And folks are wondering how to fix it.

THIS COULD BE TURNED INTO SOMETHING POSITIVE, a shining example of assessment, closing the loop, and evidence-based change.

But first, WAIT A MINUTE! Do faculty get to say "We told them what to do; if they can't comply and don't learn, it's not my fault."?  Of course not.  If students aren't learning, faculty are doing something wrong.  Lack of learning = feedback, and feedback must lead to change.

Here we have a case of an institution ignoring unambiguous feedback. The feedback is simple: the promulgation of a list of 34 things one should do on a syllabus does not produce the uniform inclusion of the small handful of actually really important things to include on a syllabus.  That's it; that's what the evidence tells you.  It doesn't tell you faculty are bad; it tells you that this method of changing what syllabi look like was ineffective.

Never mind that any good teacher knows that you cannot motivate change with a list of 34 fixes.

The correct response? Close the loop: listen, learn, change the way syllabus guidelines are handled.

The unfortunate thing here is that folks who know (faculty) brought this immediately to the attention of the folks in charge. Faculty noted that the list was too long, its provenance ambiguous, its authority unclear, its applicability variable, its tone insulting. A solution was suggested. All this was met with, basically, a brush off -- they're just guidelines not requirements, what's the big deal?

And, it turns out, that is precisely how faculty understood them. No need for alarm.  Some adjusted their syllabi to some of the suggestions in the guidelines. But apparently, the faculty didn't all implement a few of the guidelines that really do matter (to someone).  Arrrrrrrgh.

And now for a little forward looking fantasy of what the outcome of this situation COULD be.

Since administrations and the assessment industry are apparently NOT really ready to adopt the underlying premise of assessment -- pay attention to feedback and change accordingly -- the faculty will.

From now on, only the faculty will disseminate syllabi guidelines. They will very clearly distinguish between legally mandated content, accreditation relevant functionality, college-specific custom and standards, and good pedagogical practice in general. They will invite all parties who become aware of syllabi-related mandates (or new good ideas) to communicate them to the faculty's educational policy committee for consideration for inclusion in their next semester's guidelines.

Those guidelines will explicitly articulate general goals (exactly which ones to be determined) such as syllabi are to be interesting documents that are useful to students and that permit colleagues to get a sense of what a course is about and at what level it is being taught as well as suggestions of particular features, boilerplate and examples that might be useful, and fully explained required items. They will include examples of an array of syllabi that explicitly demonstrate a variety of forms that meet their standards. And, all suggestions will be referenced where possible and requirements will be documented in terms of on what authority they are an obligation.

For assessment purposes the faculty will adapt* any externally supplied "rubrics" to their own intellectually and pedagogically defensible standards and practices and encourage our colleagues to make use of these college-specific tools in developing their syllabi.

Educators really committed to the stated goals of assessment would see in this affair an opportunity for an achievement they could boast about.  Those committed to one directional, top-down, assessor-centered, non-interactive, deaf-to-feedback approaches will see in it only faculty reluctance to get with the program. 

One lesson learned here is that institutional processes need adjustment. The amount of faculty and administrative time, emotional energy, and the augmentation of frustration and mistrust that this little thing has engendered was a phenomenal waste of precious institutional resources. Alas, accountability for THIS is unlikely ever to be reckoned.

Monday, August 9, 2010

How Academic Assessment Gets it Backwards

In a letter to the NYT about an article on radiation overdoses, George Lantos writes:

My stroke neurologists and I have decided that if treatment does not yet depend on the results, these tests should not be done outside the context of a clinical trial, no matter how beautiful and informative the images are. At our center, we have therefore not jumped on the bandwagon of routine CT perfusion tests in the setting of acute stroke, possibly sparing our patients the complications mentioned.

This raises an important, if nearly banal, point: if you don't have an action decision that depends on a piece of information, don't spend resources (or run risks) to obtain the information.

Consider, for a moment, the trend toward "assessment" in contemporary higher education. A phenomenal amount of energy (and grief) is invested to produce information that is (1) of dubious validity and (2) does not, in general, have a well articulated relationship to decisions.

Now the folks who work in the assessment industry are all about "evidence based change," but they naively expect that they can, a priori, figure out what information will be useful for this purpose.

They fetishize the idea of "closing the loop" -- bringing assessment information to bear on curriculum decisions and practices -- but they confuse the means and the ends. To show that we are really doing assessment we have to find a decision that can be based on the information that has been collected.

A much better approach (and one that would demonstrate an appreciation of basic critical thinking skills) to improving higher education would be to START by identifying opportunities for making decisions about how things are done and THEN figuring out what information would allow us to make the right decision. Such an approach would involve actually understanding both the educational process and the way educational organizations work. My impression is that it is precisely a lack of understanding and interest in these things on the part of the assessment crowd that leads them to get the whole thing backwards.

Monday, December 14, 2009

Assessment and Evaluating Student Work

It's ironic, given it's centrality, how little that's sensible and defensible has been said about the relation between grading and assessment.  To my mind, it's a lost opportunity to offer constructive criticism of grading in general as well as a failure on the part of the assessment industry to demonstrate and convey clear thinking and to develop useful tools for teachers.

And so here is part one of working through a relationship between grading and assessing.

My students always want to know "how much does this count for" and so my syllabus always says something like

Exam 1
30%
Problem Sets
20%
Final Essay
40%
Participation
10%

When I grade any particular item -- assignment or paper -- I make use (at least implicitly) of a similar decomposition of the grade.  If it is an essay I may be evaluating the quality of the writing, the use of evidence, the structure of the argument, the use of sources, and so on.  If it is an exam, the questions can usually be separated into a finite number of groups, each "testing" a particular skill or understanding of a particular concept (but see fn 1 below).  Let's imagine a class in which five skills or concepts, J,K,L,M, and N, make up the content.  And let's imagine my graded activities from above can be described this way

Exam 1
25%J
25%K
25%L
25% critical
thinking

Problem Sets
20%J
20%K
20%L
20%M
20%N
Final Essay
20% Writing
30% J-N


30% Argument
20% Scholarly
conventions

Participation
33%  J-N


33% Staying up with
material in course

33% Poise, 
verbal skills, etc.



Now let's look at how all of the things I've graded fit together.  In the table below, the rows represent skills or learning outcomes that I want students to demonstrate.  The columns show me the evaluative tools I've used and which of these each one included.

Substance
Exam 1
Problem Sets
Final Essay
Participation
Concept J
+
+
+
+
Concept K
+
+
+
+
Concept L
+
+
+
+
Concept M
-
+
+
+
Concept N
-
+
+
+
Writing
-
-
+
-
Argument
-
-
+
-
Scholarly Conventions
-
-
+
-
Critical Thinking
+
-
+
-
Keeping Up
-
-
-
+
Verbal Skills
-
-
-
+

Next let's suppose that I graded each of these items on an A-F scale and that I've made some attempt to put on paper how I "operationalize" the grades "excellent," "good," "satisfactory," etc. I might, for example, have let students know that I consider an excellent use of concepts in the final essay to be when
Essay employs 3 or more of main concepts from the course in a manner that's appropriate to the subject at hand and that demonstrates a strong understanding of what they mean and how they can be useful.
And finally, let's assume that my program goals include concepts K and M and the school as a whole includes writing and critical thinking as goals.

I can simply take scores on concepts K from all four evaluations and M from the last three and then take the writing score from the final essay and the critical thinking scores from the first exam and the final essay and, oila, I've got my assessment.

fn 1  Two things require mention: 1) not every skill/concept that we expect to be learned is measured on every exam/exercise -- exams are samples; 2) many "items" will depend on more than one skill or concept.  More on these issues later.