xvii
So, the nature of the capabilities that educational assessments address can vary
over time and place, and patterns across peoples’ performances can vary in as many
ways as there are possible groups of people. Nevertheless, because we develop
capabilities in experiences shaped around recurring patterns in knowledge and
activity, there are discernable regularities. These regularities reflect recurring patterns in institutions and activities, and we can write books, design instruction, and
develop assessments that enable individuals to participate in these activities. In an
everyday sense, we speak of peoples’ RCAs or their competence in computer networking. We can communicate the results of assessment scores, even scores from
pools of tasks calibrated through latent variable models, and use the scores for grading, selection, or certification at large scale. But is this activity measurement as we
would see it through the lens of the system proposed by this book?
It appears to me that the conditions of existence, objectivity, and intersubjectivity
presented in this book must be investigated and evaluated in context. Because of the
multifarious constituents of every unique situation (but with “family resemblance”
similarities across situations) and the unique personal capabilities of individuals
(but with family resemblances that can emerge through experiences with similar
constituent patterns), we cannot uncritically expect educational assessments to provide measures as per the criteria presented in this book. We can, however, examine
the degree to which those conditions are approximated, over what ranges of persons
and situations, aided by coordinated task development, cognitive theory, and latent
variable modeling machinery. Further, we can identify groups and individuals for
whom their patterns in performance are so atypical as to preclude interpretation in
the modeling framework. There can be situations for which proceeding as if a targeted property exists, is measurable, and is approximated by a given latent variable
model is a satisfactory interpretive frame for most test-takers of interest; it may still
be that the performances of some individuals simply cannot be well understood
within that framework. (Does the putative property exist for such an individual?)
All of this makes sense to me if I take the system presented in this book, properties and measures, and variables in latent variable models as cognitive tools for us,
the analysts, to guide our thinking and our actions. Sometimes we can encounter or
construct situations in which the approximation is justified, and it is satisfactory to
think and act as if calibrated scores from a latent variable model were measures,
even as we remain alert for model misfit and departures from objectivity and intersubjectivity that distort targeted inferences. This book aims to offer an idealized
framework for measurement across science. It enables us engineers working in the
less-than-ideal real world that we can use to characterize the evidence we provide
with regard to its approximation as measures and thereby improve the quality of our
applications.
Foreword
So, the nature of the capabilities that educational assessments address can vary
over time and place, and patterns across peoples’ performances can vary in as many
ways as there are possible groups of people. Nevertheless, because we develop
capabilities in experiences shaped around recurring patterns in knowledge and
activity, there are discernable regularities. These regularities reflect recurring patterns in institutions and activities, and we can write books, design instruction, and
develop assessments that enable individuals to participate in these activities. In an
everyday sense, we speak of peoples’ RCAs or their competence in computer networking. We can communicate the results of assessment scores, even scores from
pools of tasks calibrated through latent variable models, and use the scores for grading, selection, or certification at large scale. But is this activity measurement as we
would see it through the lens of the system proposed by this book?
It appears to me that the conditions of existence, objectivity, and intersubjectivity
presented in this book must be investigated and evaluated in context. Because of the
multifarious constituents of every unique situation (but with “family resemblance”
similarities across situations) and the unique personal capabilities of individuals
(but with family resemblances that can emerge through experiences with similar
constituent patterns), we cannot uncritically expect educational assessments to provide measures as per the criteria presented in this book. We can, however, examine
the degree to which those conditions are approximated, over what ranges of persons
and situations, aided by coordinated task development, cognitive theory, and latent
variable modeling machinery. Further, we can identify groups and individuals for
whom their patterns in performance are so atypical as to preclude interpretation in
the modeling framework. There can be situations for which proceeding as if a targeted property exists, is measurable, and is approximated by a given latent variable
model is a satisfactory interpretive frame for most test-takers of interest; it may still
be that the performances of some individuals simply cannot be well understood
within that framework. (Does the putative property exist for such an individual?)
All of this makes sense to me if I take the system presented in this book, properties and measures, and variables in latent variable models as cognitive tools for us,
the analysts, to guide our thinking and our actions. Sometimes we can encounter or
construct situations in which the approximation is justified, and it is satisfactory to
think and act as if calibrated scores from a latent variable model were measures,
even as we remain alert for model misfit and departures from objectivity and intersubjectivity that distort targeted inferences. This book aims to offer an idealized
framework for measurement across science. It enables us engineers working in the
less-than-ideal real world that we can use to characterize the evidence we provide
with regard to its approximation as measures and thereby improve the quality of our
applications.
Foreword
