xiv
By the early 1930s, Sir Frederic Bartlett began to pry the black box open by
proposing schemas or meaningful patterns of knowledge in the world through which
people learn to understand a text, such as folktales, how faces look, or writing a
research proposal. Other research showed relationships of item difficulty (as
percent- correct in a targeted population) with texts’ syntactic complexities and
word frequencies (as obtained from pertinent corpora of reading materials). Under
these practices, a reading comprehension test might produce quite serviceable
scores for monitoring progress, conducting research, and identifying struggling
readers. The relationships of item features to population difficulties served as support for a provisional model of reading comprehension. Now total scores in a collection of similarly constructed tests will necessarily order persons. But this does
not guarantee that there exists a corresponding RCA property of persons, with a
range of values, for which each reader possesses a reading comprehension capability that is fully characterized by one of those values. The fact that different publishers’ choices of texts, nature of questions, and conditions of performance can reliably
sort test-takers differently casts doubt on objectivity and intersubjectivity. Similarly
troubling was the emerging evidence that some test items could prove differentially
difficult for test-takers from different ethnicities, genders, or language backgrounds
who performed similarly overall.
A latent variable modeling approach called “item response theory” (IRT), of
which the Rasch measurement model illustrated in this book may be considered an
instance, provided an analytic framework to make progress on these issues. For
familiar correct/incorrect test items, an IRT model gives the probability of a correct
response as a function of variables standing for a person’s capabilities and variables
standing for characteristics such as its difficulty and how well it sorts high and low
performers. Under idealized circumstances, these parameters and the form of the
model would account for all the systematic variation in a set of responses over some
collections of persons and items. The Rasch model in this book is a special case in
which, if it were true, the same comparisons in observed performance would hold in
probability for people regardless of items and for items regardless of people. Of
course, IRT models are never exactly true in real data such as from science and RCA
tests, and for reasons I will address shortly, we still find systematic variations for
item characteristics in different groups of people, and individuals whose response
patterns do not accord with the model very well at all.
IRT models have nevertheless proved unquestionably valuable as evidentiary
reasoning frameworks to improve practice such as enabling adaptive testing for
individuals and identifying items which are not functioning as intended. It is useful,
to be sure, as probability-based machinery to improve practical work. Do IRT models produce measures? Well, IRT models are now being extended to incorporate
cognitive theory that connects item features with process models for what people
have to know and do to solve problems. In some cases, we can construct items from
theory and predict how they will work. Other more detailed latent variable models
with categorical person variables instead of or in addition to continuous ranges connect even more closely with cognitive findings. The lid of the black box lifts further,
with process models beyond those that the trait and behavioral perspectives can
Foreword
By the early 1930s, Sir Frederic Bartlett began to pry the black box open by
proposing schemas or meaningful patterns of knowledge in the world through which
people learn to understand a text, such as folktales, how faces look, or writing a
research proposal. Other research showed relationships of item difficulty (as
percent- correct in a targeted population) with texts’ syntactic complexities and
word frequencies (as obtained from pertinent corpora of reading materials). Under
these practices, a reading comprehension test might produce quite serviceable
scores for monitoring progress, conducting research, and identifying struggling
readers. The relationships of item features to population difficulties served as support for a provisional model of reading comprehension. Now total scores in a collection of similarly constructed tests will necessarily order persons. But this does
not guarantee that there exists a corresponding RCA property of persons, with a
range of values, for which each reader possesses a reading comprehension capability that is fully characterized by one of those values. The fact that different publishers’ choices of texts, nature of questions, and conditions of performance can reliably
sort test-takers differently casts doubt on objectivity and intersubjectivity. Similarly
troubling was the emerging evidence that some test items could prove differentially
difficult for test-takers from different ethnicities, genders, or language backgrounds
who performed similarly overall.
A latent variable modeling approach called “item response theory” (IRT), of
which the Rasch measurement model illustrated in this book may be considered an
instance, provided an analytic framework to make progress on these issues. For
familiar correct/incorrect test items, an IRT model gives the probability of a correct
response as a function of variables standing for a person’s capabilities and variables
standing for characteristics such as its difficulty and how well it sorts high and low
performers. Under idealized circumstances, these parameters and the form of the
model would account for all the systematic variation in a set of responses over some
collections of persons and items. The Rasch model in this book is a special case in
which, if it were true, the same comparisons in observed performance would hold in
probability for people regardless of items and for items regardless of people. Of
course, IRT models are never exactly true in real data such as from science and RCA
tests, and for reasons I will address shortly, we still find systematic variations for
item characteristics in different groups of people, and individuals whose response
patterns do not accord with the model very well at all.
IRT models have nevertheless proved unquestionably valuable as evidentiary
reasoning frameworks to improve practice such as enabling adaptive testing for
individuals and identifying items which are not functioning as intended. It is useful,
to be sure, as probability-based machinery to improve practical work. Do IRT models produce measures? Well, IRT models are now being extended to incorporate
cognitive theory that connects item features with process models for what people
have to know and do to solve problems. In some cases, we can construct items from
theory and predict how they will work. Other more detailed latent variable models
with categorical person variables instead of or in addition to continuous ranges connect even more closely with cognitive findings. The lid of the black box lifts further,
with process models beyond those that the trait and behavioral perspectives can
Foreword
