Since its humble, and some would argue inauspicious, beginnings the focus of educational measurement (and psychological measurement, in general) has always been on latent traits; that is, “hidden, unobservable characteristics that cannot be measured directly,” but whose quantity or degree is “inferred from observable behaviors, test scores, or responses.” Intelligence, aptitude, reading ability, mathematics ability, and most recently, even proficiency in academic subjects like reading, mathematics, and science have been considered and treated as latent traits.
There was an understanding that we were not measuring the latent trait directly but instead were eliciting responses or behaviors which we used to make inferences about the trait of interest. We couldn’t see intelligence, but we could know intelligent behavior when we saw it. Specifically, when we saw varying degrees of it exhibited by people in a criterion group. For college admissions tests, the criterion might be the first-year GPA of currently enrolled college students classified. For large-scale state achievement tests, the criterion for reading or mathematics proficiency became the expected performance of a hypothetical group of students performing at the borderline between two achievement levels.
Our latent trait approach to measurement and testing worked well enough when the traits of interest like intelligence and aptitude were considered stable if not largely innate and immutable – or at least not easily changeable through instruction.
Latent traits as the foci and outcomes of testing became less tenable when our interest shifted from intelligence and aptitude to achievement.
Still even less tenable as the focus shifted from summative to formative.
Finally, the latent trait approach became untenable when rather than stable and largely immutable, the expectation was that the characteristics we were measuring and assessing could, should, and would be influenced positively by instruction.
And if there is a level below untenable (unconscionable?), we approached that limit when accountability stakes were attached to test outcomes and instructional improvement.
Out of The Darkness – Hidden No More
Educators being held accountable for performance on state tests now for the first time were primary stakeholders of large-scale test results. Rather than the indirect measurement of hidden, unobservable characteristics, their demand was for actionable information, a term which translates to test results that are timely, specific, and directly relevant to instruction. Even if timely could be achieved, results like scaled scores and achievement levels were never going to be specific enough or directly relevant to instruction.
In response to educators’ needs, in the 2000s we saw the explosion of standards-based grading, which for each of the state’s content categorized student performance along a progression. Then in the 2010s, the emergence of microcredentialing, awarding badges to students demonstrating mastery of specific targeted skills. Now in the 2020s, we have the amorphous blog that was competencies solidifying and taking shape through efforts like Skills for the Future.
Now you might question my clustering of those three distinct approaches but from my perspective, like peas in the same pod, they are closer to each other than any one of them is to the work I was doing with large-scale testing with latent traits and constructs like mathematics and reading proficiency. They share common characteristics such as being focused on parts rather than the whole and they eschew compensatory combinations like averages in favor of reporting at a more granular level. Both of those factors differ from large-scale testing and the goal to produce an overall produce an overall proficiency score.
As suggested throughout this post, however, the most critical difference from large-scale testing is the direct measurement of concreate, observable student skills and behaviors. I missed that distinction back in the 2010s when I was inviting Laurie Gagnon to the table in meetings with the Rhode Island Department of Education, and in the past few years listening to Laura Slover describe her work with competencies. Then last week as I was rereading the Skills Progressiondocument published as part of the Skills for The Future project, it struck me. The difference was there all along, as plain as the ruddy and somewhat pock-marked nose on my face: latent traits are dead.
Latent Traits Are Dead. Long Live Latent Traits.
Of course, latent traits are not dead. What may be dead, however, is the notion of treating concrete academic and durable skills as though they were latent traits. There are still latent traits worth measuring. Reading, mathematics, science, collaboration, communication, and critical thinking may not be among them.
At this time, I should probably acknowledge publicly that to a large extent the fault is not with our measurement science, but with ourselves. Throughout the history of educational measurement, we have been comfortable hiding a multitude of sins, or at least voluntary vagueness, under the invisibility cloak of latent traits. Back in 2010 during the first class of a psychometrics seminar, I asked the students to shout out the measurement, assessment, or psychometric terms they were familiar with and when the board was full, we set out to define them. When I asked for a definition of latent trait, a student confidently proclaimed that a latent trait was an undefined construct. Sadly, he wasn’t too far from the truth. From the mouths of babes, as they say.
Measures and Measurement Matter
We attempted to fit achievement testing into our existing measurement models and mechanisms. We assumed and forced unidimensionality, normal distributions, and interval scales onto skills when the content (and needs of the stakeholders) was directing us toward developmental progressions with non-interval categorizations. Perhaps there is a reason why our work on learning progressions has been stalled, if not floundering.
When designing, developing, and validating the next next generation of educational assessments it’s time that we remember and adhere to the principle that form follows function.
Certain Uncertainty
A note of caution.
As we shift from latent traits to concrete skills, from vague composites to clear descriptions grounded in content, and from decontextualized scored response strings of 0’s and 1’s to actual displays of student work, it will feel less like educational measurement and there will be a strong pull toward believing in certainty. But we must resist.
The only thing of which we can be certain is that when measuring and assessing individual students there will always be uncertainty. Whether ratings are made by humans or AI there will be uncertainty. There will be measurement error, sampling error, random and systematic error introduced by students and the environment.
If we accomplish little else in the next era of more concrete assessment, may we also find a way to make uncertainty and error less abstract, more understandable, and useful to informing decisions.
All Inferences Are Not Created Equal
A final note.
Just as there will still be uncertainty and error, there will still be a need for end users of assessment results to make inferences based on samples of student performance. When interpreting and using assessment results there will still be the need for causal inferences about achievement and growth, inferences related to contextual and program-related factors.
Inferences back to latent traits, however, will largely be replaced by inferences related to transfer. That is, inferences about how well the student will be able to apply the acquired skill in a different setting or under different conditions. The work of validating such inferences will be no less important, and likely even more visible than the validation of scores from traditional large-scale tests.
Again, if we accomplish little else this time around, let’s move validation (and with it utility) to the forefront.
Image by Alexander Lesnitsky from Pixabay