x
in ways large and small. This work might better be called educational assessment:
ways of getting and using information about how people are learning, what they can
do, how they think, how they might improve, or how they might fare in educational
or occupational settings. Colloquially, the phrase “educational measurement” situates assessment data in a quantitative framework to characterize the evidence that
the observations provide for score interpretations and score uses. How one goes
about doing this is methods, and we do this a lot. Curiously, far less attention is
accorded to more theoretical, more fundamental, questions. Just what kind of measurements might those scores be, if indeed they merit the term at all? What relation
do purported measures, as numerals and categories, have to attributes of people?
A conjunction of factors, I believe, led to this relative lack of interest. First,
assessment itself was familiar. Examinations have been used for more than a millennium, as for selection in civil service in imperial China and matriculation in medieval European universities in Bologna, Paris, and elsewhere. These scores had
authority by way of authorities!
Second, the measurement of quantitative properties in the physical sciences had
by the turn of the twentieth century earned authority the hard way: by theory, evidence, argumentation, instrumentation, and demonstrated coherence within a web
of scientific and practical phenomena. Such a measure holds value not simply
because it produces numbers, but because it places a trace of a unique event,
observed at a particular time and place in a certain way, into a network of regular
relationships among objects and events that holds meaning across times, places,
and people. Being able to calibrate local measures to a common metric is a hallmark
of physical measurement. There is typically a dialectic between improved theory
and improved instrumentation: increasingly efficacious theories of properties and
the processes by which they bring about effects under instrumentation (i.e., “opening the black box”). To the shared benefit of science and commerce, national and
international institutions such as the International Bureau of Weights and Measures
standardize units and vocabularies. As two running examples, this book uses length,
the canonical case of classical quantitative measurement, and the more nuanced and
therefore illuminating case of temperature.
Third, psychologists sought to extend this quantitative measurement frame to the
psychological and social realm, or psychosocial sciences in the terminology of this
book. Psychophysicists began studying sensory perception and acuity through the
lens of measurable human attributes in the mid-1800s, and it is here that philosophical controversies about the nature of measurement in human sciences surfaced.
Early psychometricians, such as Charles Spearman and Louis Terman, adapted their
methods to data from the emerging standardized tests of educational and psychological constructs such as intelligence, verbal aptitude, and reading comprehension
ability (RCA), the running example of this book. My oversimplification: (1) a test
was crafted to educe a trait thought to exist as a property of persons; (2) persons’
interactions with the test situations are coded to produce numbers that are taken to
indicate more or less of that property; and (3) the numbers are taken as measures of
values of said property for each person. This chain of reasoning is replete with
terms (italicized above) that could be defined in multiple ways (which have practical
Foreword
in ways large and small. This work might better be called educational assessment:
ways of getting and using information about how people are learning, what they can
do, how they think, how they might improve, or how they might fare in educational
or occupational settings. Colloquially, the phrase “educational measurement” situates assessment data in a quantitative framework to characterize the evidence that
the observations provide for score interpretations and score uses. How one goes
about doing this is methods, and we do this a lot. Curiously, far less attention is
accorded to more theoretical, more fundamental, questions. Just what kind of measurements might those scores be, if indeed they merit the term at all? What relation
do purported measures, as numerals and categories, have to attributes of people?
A conjunction of factors, I believe, led to this relative lack of interest. First,
assessment itself was familiar. Examinations have been used for more than a millennium, as for selection in civil service in imperial China and matriculation in medieval European universities in Bologna, Paris, and elsewhere. These scores had
authority by way of authorities!
Second, the measurement of quantitative properties in the physical sciences had
by the turn of the twentieth century earned authority the hard way: by theory, evidence, argumentation, instrumentation, and demonstrated coherence within a web
of scientific and practical phenomena. Such a measure holds value not simply
because it produces numbers, but because it places a trace of a unique event,
observed at a particular time and place in a certain way, into a network of regular
relationships among objects and events that holds meaning across times, places,
and people. Being able to calibrate local measures to a common metric is a hallmark
of physical measurement. There is typically a dialectic between improved theory
and improved instrumentation: increasingly efficacious theories of properties and
the processes by which they bring about effects under instrumentation (i.e., “opening the black box”). To the shared benefit of science and commerce, national and
international institutions such as the International Bureau of Weights and Measures
standardize units and vocabularies. As two running examples, this book uses length,
the canonical case of classical quantitative measurement, and the more nuanced and
therefore illuminating case of temperature.
Third, psychologists sought to extend this quantitative measurement frame to the
psychological and social realm, or psychosocial sciences in the terminology of this
book. Psychophysicists began studying sensory perception and acuity through the
lens of measurable human attributes in the mid-1800s, and it is here that philosophical controversies about the nature of measurement in human sciences surfaced.
Early psychometricians, such as Charles Spearman and Louis Terman, adapted their
methods to data from the emerging standardized tests of educational and psychological constructs such as intelligence, verbal aptitude, and reading comprehension
ability (RCA), the running example of this book. My oversimplification: (1) a test
was crafted to educe a trait thought to exist as a property of persons; (2) persons’
interactions with the test situations are coded to produce numbers that are taken to
indicate more or less of that property; and (3) the numbers are taken as measures of
values of said property for each person. This chain of reasoning is replete with
terms (italicized above) that could be defined in multiple ways (which have practical
Foreword
