Showing posts with label assessment. Show all posts
Showing posts with label assessment. Show all posts

Saturday, April 7, 2012

What If Administrator Pay Were Tied to Student Learning Outcomes

The recent negotiation in Chicago ("Performance Pay for College Faculty") of a tie between student performance and college instructor pay brought this accolade from an administrator:  it gets faculty "to take a financial stake in student success."

It got me wondering why we don't hear more about directly tying administrator pay to student success.  If we did, I'll bet the students would have a lot more success.  At least, that's what the data released to the public (and Board of Trustees) would show.  There'd be far less of a crisis in higher education.

Thought experiment. What would  happen if we were to tie administrator pay to student success -- much the way corporate CEOs have their pay packages designed -- especially administrators of large multi-campus systems.

Prediction 1.  The immediate response to the very proposal would be "oh, no, you can't do that because we do not have the same kind of authority to hire and fire and reward and punish that a corporate CEO has."  But think about this...
  1. Private sector management has a lot less flexibility than those looking in from the outside think.  Almost all of the organizational impediments to simple, rational management are endemic to all organizations.
  2. Leadership is not primarily about picking the members of your team. It's about what you manage to get the team you have to accomplish.
  3. Educational administrators do not start the job ignorant of how these educational institutions work. It is tremendously disingenuous to say "if only I had a different set of tools."  People who do not think they can manage with the tools available and within the culture as it exists should not take these jobs in the first place.
  4. This, it turns out, is what some people mean when they say that schools should be run like a business. The first impulse of unsuccessful leaders is to blame the led. The second one is to engage in organizational sector envy: "if I had the tools they have over in X industry...."  What this ignores is the obvious evidence that others DO succeed in your industry with your tools.  And plenty of leaders "over there" fail too.  It is not the tools' fault.
Prediction 2.  Learning would be redefined in terms of things produced by inputs administrators had more control over.  And resources would flow in that direction too.

Prediction 3. Administrators would get panicky when they looked at the rubrics in the assessment plans they exhort faculty to participate in and that are included in reports they have signed off on for accreditation agencies.  They'd suddenly start hearing the critics who raise questions about methodologies.  They would start to demand that smart ideas should drive the process and that computer systems should accommodate good ideas rather than being a reason for implementing bad ones.

Prediction 4. In some cases it would motivate individuals to start really thinking "will this promote real learning for students" each time they make a decision.  And they'll look carefully at all that assessment data they've had the faculty produce and mutter, "damned if I know."

Prediction 5. Someone will argue that the question is moot because administrators are already held responsible for institutional learning outcomes.   Someone else will say "Plus ça change, plus c'est la même chose."

Better Teaching Through a Financial Stake in the Outcome

In an Inside Higher Ed article this week ("Performance Pay for College Faculty") K Basu and P Fain describe how the new contract signed between City Colleges of Chicago and a union representing 459 adult education instructors links pay raises to student outcomes.

Administrators lauded the move in part because it gets faculty "to take a financial stake in student success." The details of the plan are not clear from the article, but the basic framework is to use student testing to determine annual bonus pay for groups of instructors working in various areas. That is, in this particular plan it does not sound like the incentive pay is at the level of individual instructors.

Still, should the rest of higher education be paying attention? Adult education at CCC is, after all, a markedly different beast than full time liberal arts institutions or 4 year state schools or research universities. One reason we should because it's precisely the tendency to elide institutional differences that is one of the hallmarks of the style of thought endemic among some higher education "reformers." Those who think it's a good idea for adult education institutions are likely to champion it elsewhere.

But most germane for the subject of this blog is the question of what data would inform such pay for performance decisions when they are proposed for other parts of American higher education. Likely it will be something that grows out of what we now know as learning assessment. I ask the reader: given what you have seen of assessment of learning outcomes in your college, how do you feel about having decisions about your pay check based upon it?

But, your opinion aside, there are several fundamental questions here. One is whether you become a more effective teacher by having a financial stake in the outcome. The industry where this incentive logic has been most extensively deployed is probably the financial services industry, especially investment banking.  How has that worked for society?  It would be easy to cook up scary stories of how this could distort the education process, but that's not even necessary to debunk the idea.  The amounts at play in the teacher pay realm are so small that one can barely imagine even a nudge effect on how people approach their work.

But what about the data?  Consider the prospect of assessment as we know it as input to ANY decision process, let alone personnel decisions.  Anyone who has spent any time at all looking at how assessment is implemented knows that the error bars on any datum emerging from it dwarf the underlying measurement. The conceptual framework is thrown together on the basis of dubious theoretical model of teaching and learning and forced collaboration between instructors and assessment professionals.  The process sacrifices methodological rigor in the name of pragmatism, a culture of presentation (vis a vis accreditation agencies), and the tail of design limitations of software systems that wags the dog of pedagogy and common sense.  At every step of the process information is lost and distorted. But it seems that the more Byzantine that process is, the more its champions think they have scientific fact as product.

It could well be that the arrangement agreed to in Chicago will lead to instructors talking to one another about teaching, coordinating their classroom practices, and all sorts of other things that might improve the achievements of their students.  But it will likely be a rather indirect effect via the social organization of teachers (if I understood the article, the good thing about the Chicago plan is that it rewards entire categories of instructors for the aggregate improvement).  To sell it at the level of individual incentive is silly and misleading.  And, if we think more broadly about higher education, the notion that you can take the kinds of discrimination you get from extremely fuzzy data and multiply it by tiny amounts of money to produce positive change at the level of the individual instructor is probably best called bad management 101.

Sunday, September 25, 2011

Rubrics, Disenchantment, and Analysis I

There is a tendency, in certain precincts in, and around, higher education, to fethishize rubrics.  One gets the impression at conferences and from consultants that arranging something in rows and columns with a few numbers around the edges will call forth the spirit of rational measurement, science even, to descend upon the task at hand.  That said, one can acknowledge the heuristic value of rubrics without succumbing to a belief in their magic.  Indeed, the critical examination of almost any of the higher education rubrics in current circulation will quickly disenchant, but one need not abandon all hope: if assessment is "here to stay," as some say, it need not be the intellectual train wreck its regional and national champions sometimes seem inclined to produce.

Consider this single item from a rubric used to assess a general education goal in gender:


As is typical of rubric cell content, each of these is "multi-barrelled" -- that is, the description in each cell is asking more than one question at a time. It's not unlike a survey in which respondents are asked, "Are you conservative and in favor of ending welfare?"  It's a methodological no-no, and, in general, it defeats the very idea of dis-aggregation (i.e., "what makes up an A?") that a rubric is meant to provide.

In addition, rubrics when they are presented like this are notoriously hard to read. That's not just an aesthetic issue -- failure to communicate effectively leads to misuse of the rubrik (measurement error) and reduces the likelihood of effective constructive critique.

Here is the same information presented in a manner that's more methodologically sound and more intellectually legible:

At the risk of getting ahead of ourselves, there IS a serious problem when these rank ordered categories are used as scores that can be added up and averaged, but we'll save that for another discussion.  Too, there is the issue of operationalization -- what does "deep" mean, after all, and how do you distinguish it from not so deep?  But this too is for another day.

Let's, for the sake of argument, assume that each of these judgments can be made reliably by competent judges. All told, 4 separate judgments are to be made and each has 3 values. If these knowledges and skills are, in fact, independent (if not, a whole different can of worms), then there are 3 x 3 x 3 x 3 = 81 combinations of ratings possible. Each of these 81 possible assessments is eventually mapped on to1 of 4 ratings. Four combinations are specified, but the other 77 possibilities are not:

Now let us make an (probably invalid) assumption: that each of THESE scores is worth 1, 2 or 3 "points" and then let's calculate the distance between each of the four scores. We use standard Euclidean distance – r=sqrt(x2 + y2) with the categories being: Mastery = 3 3 3 3, Practiced = 2 2 2 3, Introduced = 2 2 2 2, Benchmark = 1 1 1 1


So, how do these categories spread out along the dimension we are measuring here? Mastery, Introduced, and Benchmark are nicely spaced, 2 units apart (and M to B at 4 units). But then we try to fit P in. It's 1.7 units from Mastery and 2.2 from Benchmark, but it's also 1 unit from Introduced. To represent these distances we have to locate it off to the side.

This little exercise suggests that this line of the rubrik is measuring two dimensions.

This should provoke us into thinking about what dimensions of learning are being mixed together in this measurement operation.

It is conventional in this sort of exercise to try to characterize the dimensions in which the items are spread out. Looking back at how we defined the categories we speculate that one dimension might have to do with skill (analysis) and the other knowledge. But Mastery and Practiced were on the same level on analysis. What do we do?


It turns out that the orientation of a diagram like this is arbitrary -- all it is showing us is relative distance. And so we can rotate it like this to show how our assessment categories for this goal relate to one another.

Now you may ask what was the point of this exercise?  First, if the point of assessment is to get teachers to think about teaching and learning, and to do so in a manner that applies the same sort of critical thinking skills that we think are important for students to acquire then a careful critique of our assessment methods is absolutely necessary.

Second, this little bit of quick and dirty analysis of a single rubric might actually help people design better rubrics AND to assess the quality of existing rubrics (there's lots more to worry about on these issues, but that's for another time).  Maybe, for example, we might conceptualize "introduce" to include knowledge but not skill or vice versa?  Maybe we'd think about whether the skill (analysis) is something that should cross GE categories and be expressed in common language.  And so on.

Third, this is a first step toward showing why it makes very little sense to take the scores produced by using rubrics like this and then adding them up and averaging them out in order to assess learning.  That will be the focus of a subsequent post.

Sunday, August 23, 2009

Let's Take It Seriously

Let's take assessment and accountability seriously AS AN INSTITUTION. There is a tendency to equate assessment with measuring what professors do to/with students. The buzz word is "accountability" and there's this unspoken assumption that the locus of lack of accountability in higher education is the faculty. I think that assumption is wrong.

We should broaden the concept of assessment to the whole institution. Course instructors get feedback on an almost daily basis -- students do or don't show up for class; instructors face 20 to 100 faces projecting boredom or engagement several times per week; students write papers and exams that speak volumes about whether they are learning anything; advisees tell faculty about how good their colleagues are. By contrast, the rest of the institution has little, if any, opportunity for feedback. It's important: one substandard administrative act can affect the entire faculty, so even small things can have a big negative effect on learning outcomes.

In the name of accountability throughout the institution I propose something simple, but concrete: every form or memo should have a "feedback button" on it. Clicking on this button will allow "users" anonymously to offer suggestions or criticism. These should be recorded in a blog format -- that is, they accumulate and are open to view. At the end of each year, the accountable officer would be required in her or his annual report to tally these comments and respond to them, indicating what was learned, what changes have been made or why changes were not made.

The important component of this is that the comments are PUBLIC so that constituents can see what others are saying. Each "user" can see whether her ideas are commonly held or idiosyncratic and the community can know what kind of feedback an office is receiving and judge its responsiveness accordingly.

Why anonymous? This is feedback, not evaluation. This information cannot be used to penalize or injure anyone. The office has opportunity to respond either immediately or in an annual report. Crank comments will be weeded out by sheer numbers and users who will contradict them. In the other direction, it is clear that honest feedback can be compromised by concerns about retribution, formal or informal. Further analysis along these lines would further support the idea that comments should be (at least optionally) anonymous.

We should note that we already do all of this in principle -- many offices around campus have some version of a "suggestion box." What is missing is (1) systematic and consistent implementation so that users get accustomed to the process of providing feedback, and (2) a protocol for using the feedback to enrich the community knowledge pool and to build it into an actual accountability structure.

The last paragraph makes the connection to a sociology of information. Information asymmetries (as when the recipient knows what the aggregate opinion is, but the "public" does not) and the atomization of polities (this is what happens when opinion collection is done in a way that minimizes interactions among the opinion holders -- cf. Walmart not wanting employees to discuss working conditions -- preventing the formation of open, collective knowledge*) are a genuine obstacle to organizational improvement. Many, many private organizations have learned this; it's not entirely surprising that colleges and universities are the last to get on board.

* as opposed, say, to things that might be called "open secrets"

Wednesday, August 5, 2009

Validity and Such

An AACU blogpost referred me to the National Institute for Learning Outcomes Assessment website which referred me to an ETS website about the Measure of Academic Proficiency and Progress (MAPP) where I would be able to read an article titled "Validity of the Measure of academic Proficiency and Progress ."

And here's the upshot of that article: The MAPP is basically the same as the test it replaced and research on that test showed
...that the higher scores of juniors and seniors could be explained almost entirely by their completion of more of the core curriculum, and that completion of advanced courses beyond the core curriculum had relatively little impact on Academic Profile scores. An earlier study (ETS, 1990) showed that Academic Profile scores increased as grade point average, class level and amount of core curriculum completed increased.
In other words, the test is a good measure of whether students took more GenEd courses. And we suppose that in GenEd courses students are acquiring GenEd skills. And so these tests are measures of the GenEd skills we want students to learn.

A tad circular? What exactly is the information value added by this test?