Monday, December 14, 2009

Assessment and Evaluating Student Work

It's ironic, given it's centrality, how little that's sensible and defensible has been said about the relation between grading and assessment.  To my mind, it's a lost opportunity to offer constructive criticism of grading in general as well as a failure on the part of the assessment industry to demonstrate and convey clear thinking and to develop useful tools for teachers.

And so here is part one of working through a relationship between grading and assessing.

My students always want to know "how much does this count for" and so my syllabus always says something like

Exam 1
30%
Problem Sets
20%
Final Essay
40%
Participation
10%

When I grade any particular item -- assignment or paper -- I make use (at least implicitly) of a similar decomposition of the grade.  If it is an essay I may be evaluating the quality of the writing, the use of evidence, the structure of the argument, the use of sources, and so on.  If it is an exam, the questions can usually be separated into a finite number of groups, each "testing" a particular skill or understanding of a particular concept (but see fn 1 below).  Let's imagine a class in which five skills or concepts, J,K,L,M, and N, make up the content.  And let's imagine my graded activities from above can be described this way

Exam 1
25%J
25%K
25%L
25% critical
thinking

Problem Sets
20%J
20%K
20%L
20%M
20%N
Final Essay
20% Writing
30% J-N


30% Argument
20% Scholarly
conventions

Participation
33%  J-N


33% Staying up with
material in course

33% Poise, 
verbal skills, etc.



Now let's look at how all of the things I've graded fit together.  In the table below, the rows represent skills or learning outcomes that I want students to demonstrate.  The columns show me the evaluative tools I've used and which of these each one included.

Substance
Exam 1
Problem Sets
Final Essay
Participation
Concept J
+
+
+
+
Concept K
+
+
+
+
Concept L
+
+
+
+
Concept M
-
+
+
+
Concept N
-
+
+
+
Writing
-
-
+
-
Argument
-
-
+
-
Scholarly Conventions
-
-
+
-
Critical Thinking
+
-
+
-
Keeping Up
-
-
-
+
Verbal Skills
-
-
-
+

Next let's suppose that I graded each of these items on an A-F scale and that I've made some attempt to put on paper how I "operationalize" the grades "excellent," "good," "satisfactory," etc. I might, for example, have let students know that I consider an excellent use of concepts in the final essay to be when
Essay employs 3 or more of main concepts from the course in a manner that's appropriate to the subject at hand and that demonstrates a strong understanding of what they mean and how they can be useful.
And finally, let's assume that my program goals include concepts K and M and the school as a whole includes writing and critical thinking as goals.

I can simply take scores on concepts K from all four evaluations and M from the last three and then take the writing score from the final essay and the critical thinking scores from the first exam and the final essay and, oila, I've got my assessment.

fn 1  Two things require mention: 1) not every skill/concept that we expect to be learned is measured on every exam/exercise -- exams are samples; 2) many "items" will depend on more than one skill or concept.  More on these issues later.

Wednesday, November 25, 2009

Responding to Student Writing

In a piece called "ABOUT RESPONDING TO STUDENT WRITING," Peter Elbow writes:
The fact is there is no best way to respond to student writing. The right comment is the one that will help this student on this topic on this draft at this point in the semester -- given her character and experience.My best chance for figuring out what is going on for any particular student at any given point depends on figuring out what was going on for her as she was writing.
I found this document on our schools website under "Teaching Resources" on the Provost Office page. It sounds like pretty good advice.

It also sounds different from the advice I find on another page of the institution's website. On that page I find a "rubric" for assessing student learning in essays. It gives me six categories (overall impression, argument, evidence, counter evidence, sources, citations) and wordy descriptions of different levels of achievement in each. It's pretty unclear from the document how it is intended to be used, but basically, it's a grading scale.

Here's my question: which kind of teaching students to write does my boss want me to use?

Friday, November 6, 2009

Spellings' Flawed Metaphor

An interesting article and even more interesting responses in Inside Higher Ed from a few years back: "The Flawed Metaphor of the Spellings Summit"

Monday, November 2, 2009

Six Things to Beware of, Grasshopper...

A lot of such talk as there is about innovation and change in higher education these days shows up in the general orbit of assessment. Having recently listened to or read a lot of material from assessment experts, I jotted down a few cautions.

Beware the tyranny of software. I like software. I write computer programs. But, at the risk of sounding Asimovian, software has to serve education, not the other way round. For the last ten years we have repeatedly adjusted the way we educate to the needs of the software ("Banner won't let you do that..."). Banner and its ilk are just hammers. No right-thinking carpenter changes the way she builds a house because of hammer limitations.

Beware the fetishization of uniformity. Healthy ecosystems, organizations, relationships, families, and individuals entertain a healthy dialectical tension between sameness and difference, uniformity and irregularity, standardization and improvisation. Whether you are trying to be a virus that can outwit immune systems (or an immune system that can shut down a virus), or building a firm that can weather economic ups and downs, or running a college that produces excellent graduates from a stunning variation of inputs, the key is to cultivate order and chaos simultaneously.

Beware projection and other forms of X-o-centrism. We are all subject to our own version of Saul Steinberg's classic "New Yorker's view of the World." What works for me, or in one course, or in my department, or in our division, or in one school I know about, is surely good for you. Whether it's a analogy or metaphor, an algorithm, a social form, or a paradigm, or a homeomorphism, context and local history matter.

Beware foolish numberers and their arrogant misquotations of Lord Kelvin -- if you can't measure it, it does not exist -- and mindless adherence to quantification as an end in itself.

Beware saviors, those who in the face of skepticism and critique fancy themselves the new Galileo or who too readily imagine they are members of a new Vienna Secession or Salon des Refusés.

Beware assurances that complex things can be done with little effort or in far less time than you think. Most things that are easy and simple and beneficial have already been done. Things like the valid measurement of educational outcomes are not simple. Getting it right takes time and effort.

Most of all, beware a movement that cannot apply its own techniques to itself.

When the Blind Meet the Lost

Assessment has landed where it has because the "movement" is driven, far beyond our individual institutions' halls, by a political agenda and small minds who have seized on an entrepreneurial opportunity and attached themselves to it.   In a democratic society, that politcal agenda deserves a free and open debate.  Unfortunately, many of the individuals who have attached themselves to it, are either unaware of its terms or incapable (or afraid) of engaging in such a debate.  Unfortunately for those who are behind the movement, many of their foot soldiers are an embarrassment and either they themselves are not competent to realize this or they are too ideologically blinded to care.

Perhaps the most telling characteristic of the "assessment movement" is its failure to live up to its own standards: there is no culture of accountability and measurement and assessment in the assessment community. It would not be the first movement (in education or elsewhere) to suffer from this shortcoming.

Friday, October 30, 2009

The Fetishization of Rubrics I

The one thing you see over and over and over in the assessment literature is the "rubric."  Never mind, for now, the history of the concept -- that's an interesting story but it's for another time.

For now, just a quick note.  A rubric is basically a two dimensional structure, a table, a matrix.  The rows represent categories or observable or measurable phenomena (such as, for grading an essay, "statement of topic," "grammar," "argument," and "conclusion") and the columns represent levels of achievement (e.g., "elementary," "intermediate," "advanced").  The cells of the table then contain a description of the level of "grammar" that would constitute different levels of performance.

A rubric is, we could say, just a series of scales that use the same values with something like the operationalization of each value specified.

Rubrics are, in other words, nothing new.  Why then, our first question must be, do assessment fanatics act as if rubrics are new, something they have discovered and delivered to higher education?

I would submit that the answer is ignorance and naivete.  They just don't know.

A second question is why their rubrics are so often so unsophisticated.  Most rubrics you find on assessment websites, for example, suggest no appreciation for something as elementary as the difference between ordinal, interval, and ratio measurements.  Take this one, which is a meta-rubric (an assessment rubric for rating efforts at assessment using rubrics). (Source: WASC)

Criterion

Initial

Emerging

Developed

Highly Developed
Comprehensive List
Assessable Outcomes
Alignment
Assessment Planning
The Student Experience

Looks orderly enough, eh? Let's examine what's in one of the boxes. Here's the text for "Assessment Planning" at the "Developed" level:
The program has a reasonable, multi-year assessment plan that identifies when each outcome will be assessed. The plan may explicitly include analysis and implementation of improvements.
It looks like we need another rubric because we've got lots going on here:
  1. What makes a "reasonable, multi-year plan"?
  2. Mainly what we need here are dates : when will each outcome be assessed.
  3. How should the assessor rate the "may-ness" of analysis and implementation? Apparently these do not make the plan better or worse since they may or may not be present.
Our next analytical step might be to look at what varies between the different levels of "Assessment Planning" but first let's ask what conceptual model lies behind this approach?  It's very much that of developmental studies, especially psychology.   The columns are, thus, stages of development.  In psychology or child development the columns have some integrity in the natural stages a person goes through.  Dimensions may be independent in terms of measurement but highly correlated (typically with chronology) and so "stages" emerge naturally from the data.

In the case of assessment, though, these are a priori categories made up by small minds who like to put things in boxes.  And the analogy they are making when they make them up is very much to child development.  An assessment rubric is a grown up version of kindergarten report cards.

Friday, August 28, 2009

Why Do We Need a Faculty Assessment Committee?

Any time we create a committee we should stop and ask why.  The baseline for answering that question should be the world (or institution) without the committee.  How was it?  How would it be?

When I think about that in this case, here's what I come up with.  With no practicing assessment committee (it was appointed but never met last year):
  1. Faculty have felt little opportunity for real input into assessment
  2. The process has in fact, over the years, been dominated by non-faculty and non-academics.
  3. Many faculty members are unimpressed with the process. Substantive missteps have been frequent.  Faculty members' assessments of assessment span the range from feeling insulted by the unprofessional and intellectually demeaning tone with which assessment has frequently been conveyed to serious criticism of the validity of the methods used in assessment and real concern about how it is consistently ignored or dismissed.  And much in between.
[I suspect that from the "other side" it looks like this
  1. Faculty have been slow to adopt a culture and practice of assessment
  2. Our job is to get the institution to comply with WASC enough to get us re-certified]
So how to make the world different WITH an assessment committee?  If I were an administrator, I'd think that the committee could help me to bring the faculty along.   I could co-opt them as fellow champions of assessment as currently practiced and they'd be vanguards of the movement.

Uh, I don't think so.  The problem with assessment is not lack of faculty buy-in.  Let's repeat that: THE PROBLEM WITH ASSESSMENT IS NOT FACULTY BUY-IN.  The problem with assessment is (are):
  • its methods are methodologically dubious
  • its logic model (observation>analysis>change) is vague, rarely made explicit, and more wishful thinking than realistic
  • it dishonestly or naively hides its political values behind a veil of "objective measurement"
  • it is dominated by self-serving educational entrepreneurs who live off, not for, assessment
  • it is evangelized in the absence of hard thinking about institutional inputs and outputs, the very things it purports to be sensitive to
  • it enters the academy as a fait accompli, more based on conviction and belief than theory, analysis, and argument, and exempts itself from the critical examination and culture of evidence that it champions
So, what can an assessment committee do if even some of the above is in fact that case?  Mainly, I think, hold assessment accountable to normal standards of intellectual integrity and professionalism.  If we do that, I predict, there would be changes in how assessment is implemented, changes that would allow the process to capitalize on its virtues and avoid some of its vices.  And in reaction to THAT you would get more buy in.  The giant flaw in how it's been handled so far is that assessment is blind to closing its own loop.  When faculty don't fall in line, it's not necessarily because they are resistant to change, unwilling to give up their comfortable sinecures, or too arrogant to think about students.  Sometimes its because they have looked at something, and, smart people that they are, found it wanting.

It may even be that the resistance to change and feedback, the comfortable sinecures, and the arrogance that deflects all criticism may lie in the assessment industry itself.  The rest is, as they say, projection.