Danger! Steep Grades Ahead

Grading – it’s as easy as ABC (plus D and E/F). 

Except, of course, grading isn’t easy at all.  Far from it. 

As the informational blurb for Guskey and Brookhart’s 2019 book What We Know About Grading states, “grading is one of the most hotly debated topics in education, and grading practices are largely based on tradition, instinct, or personal history or philosophy.” A 2023 Harvard Graduate School of Education piece by Lori Hough, The Problem with Grading, provide examples of the idiosyncrasies of individual teachers applying their various traditions, instincts, personal histories and philosophies to assigning grades to students, often students within the same school or taking the same course. 

So, grading’s a problem. What was the solution? 

Without hyperbole, one can argue that large-scale K-12 testing, the lifeblood for so many of us reading this blog, owes its very existence to those idiosyncrasies in grading. From Horace Mann to Arne Duncan, it was the lack of standardization in teachers’ grading that lit the spark for some form of standardized testing. For nearly two centuries now, grades were the problem and tests were the solution. The irony that standardized testing has had countless deleterious effects on grading (and instruction) while seemingly having had no impact on reducing those idiosyncrasies is not lost on anyone, but neither has it slowed down the testing train.

As those of us involved in testing kept our distance from the messiness of schools and schooling, we could sit comfortably above the fray as inconsistencies between test scores and grades only fueled the fires which fanned the flames of calls for more and better testing. Now, however, the call is to better integrate assessment and instruction, to connect the dots on the curriculum-instruction-assessment triangle, to close the loop on the Formative Assessment Cycle. 

It’s going to get messy

History tells us that integrating instruction and assessment will not be easy. 

Our recent attempt to stick a toe into the classroom waters by weighing in with a technical “measurement” solution to a seemingly straightforward problem with grading practices (i.e., assigning ‘0’ scores) serves as fair warning of just how challenging it will be to bring our work on assessment into the classroom in the service of learning. 

Compared to understanding how students learn and real-time decisions made on the fly during instruction, grading is relatively low-hanging fruit. However, …

The Science of Grading? Or …Not Everything That Counts Can Be Counted

Guskey and Brookhart attempt to synthesize a century of research to identify “what works” and ‘what doesn’t” with regarding to grading. Sadly, however, it appears that there is no “Science of Grading” analogous to the “Science of Reading” that we can turn or point to with confidence. I did come across a 1953 Science Education article by Otto JM Smith, an engineer and professor from UC Berkeley, titled The science of grading, which addresses many of the common issues associated with trying to combine student work across a semester or academic year to produce a composite score. Sadly, those issues are still common and not at all surprising. TL: DR summary: “averages suck.”

Alas, three-quarters of a century later K-12 teachers and higher education professors still rely heavily on computing averages to produce a composite score, and we have no uppercase or lowercase science of grading. 

To be honest and only a tiny bit cynical, if I were to sign next month’s social security check over to one of those fancy new prediction markets, I would wager on seeing any one of the following articles, books, webinars, podcasts, or movements appear before the seminal work on the science of grading:

  • The Tao of Grading
  • The Zen of Grading
  • The Art of Grading
  • The Politics of Grading
  • The Sociocultural Affect and Effects of Grading
  • Gimme an “A”: A Comprehensive Guide to Balanced Grading Systems

Why? 

Because…

There are more things in heaven and earth, Horatio…

The bottom line is that student grading has always been about so much more than finding the best way to produce a composite score or rating from the available samples of student work. At its best, grading is controlled chaos in the technical sense of the term: “positioning a system at that edge where it’s flexible enough to respond to surprises but structured enough to avoid collapsing into pure randomness.” 

Even when we eliminate the more egregious practices associated with grading (e.g., adding or subtracting points for non-academic activities or behaviors) and support teachers with assessments better aligned to critical standards there will still be the need for flexibility in grading. 

Situations centered around humans, particularly dynamic situations involving teaching, learning, and complex young humans still growing in so many ways needs such flexibility. Anyone who has sat on either side of a gradebook knows that to be true. What happens when you remove flexibility from the system? We have already witnessed the unmitigated disaster that resulted from misguided attempts to wring the flexibility out of grading by turning it over to first-generation machines in the form of electronic gradebooks. 

With all due respect to John Greenleaf Whittier, there are few sadder words than those of a frustrated teacher stating, “There’s nothing I can do, it’s already in the gradebook.”

A byproduct of the inflexibility of the electronic gradebook are issues associated with the practice of assigning ‘0’ scores, specifically the disproportionate and negative impact that such scores have on student grades and students themselves. It was on this issue that several of my assessment brethren decided to weigh in recently and over the past few years. 

Enter The Tester (rhymes with jester) – A False First Step

With few exceptions, voices from the assessment community have been opposed to assigning ‘0’ scores and in favor of a lower limit such as ‘50’ on tests and other student work scored out of a possible 100 points. Arguably, a lower bound of ‘50’ is a sound technical recommendation given the issue of the inflexibility of electronic gradebook raised above. And policies establishing a floor of ‘50’ on one-shot high stakes tests such as a mid-year or final exam were in place when I was teaching back in the 1980s. 

As noted, however, although certainly there are technical aspects to assessing student performance, grading is not fundamentally a technical problem. 

In classes where already it is common practice for the difficulty of tests and assignments to be calibrated with student achievement and expectations for student performance, ‘0’ scores are rarely a reflection of actual student performance. More often than not they reflect students not participating in the instruction/assessment process for one reason or another. 

When, as is our wont, we viewed the grading issue from one dimension (i.e., the electronic gradebook) we took away the flexibility needed to evaluate the various reasons for students not participating and to act accordingly. Replacing an inflexible gradebook with an equally inflexible grading policy fails to appreciate the complexity of the instructional environment and student-teacher interactions. Worst case, that lack of understanding leads to the viral 2018 social media post from a middle school teaching claiming she was fired for refusing to give 50s to students who did not turn in a project. True, not true, or most likely somewhere in between, her post reemerges each school year and sparks passionate online debate about the complicated world of grading policies and practices. 

And grading is only the uppermost tip of a very large iceberg.

Humble and Brave in A New World

It is humbly and bravely that we must step into the complex and multidimensional world of schools filled with dynamic interactions among teachers and students. 

Humbly, because it is not our world and we have little understanding of it. As I wrote in the inaugural post for this blog back in 2015, “it is much easier for teachers to understand psychometrics than for psychometricians to understand teaching.” That is, what teachers need to understand about our world pales in comparison to what we now need to understand about theirs. 

We have devoted a lot of effort toward improving educators’ assessment literacy. Now it is time to increase our understanding of the educational part of educational measurement or educational assessment. 

Bravely, because the complexity of what we are being asked to do to support instruction far exceeds the task that we have performed routinely and for so many years with our large-scale tests; that is, estimating student achievement on a fixed set of content at a fixed point in time under fairly standardized and controlled conditions. In terms of complexity, the information that we generate is much more similar to collecting attendance data than it is to the information needed to support instruction. 

The learning curve for integrating assessment with instruction will be steep and the road ahead challenging. 

It’s time to get started. 

Image by Clker-Free-Vector-Images from Pixabay

Published by Charlie DePascale

Charlie DePascale is an educational consultant specializing in the area of large-scale educational assessment. When absolutely necessary, he is a psychometrician. The ideas expressed in these posts are his (at least at the time they were written), and are not intended to reflect the views of any organizations with which he is affiliated personally or professionally..