90
ity is a single property of a test. Second, Messick reframed validity from being a
claim about the true state of affairs (“a test measures what it claims to measure”) to
being a claim about the present state of available evidence, as judged by a particular
community—that is, from an ontological claim to an epistemic claim: “validity is an
integrated evaluative judgment of the degree to which empirical evidence and theoretical rationales support the adequacy and appropriateness of inferences and actions
based on test scores” (Messick, 1989: p. 13). The idea of distinct types of validity
(e.g., criterion, content, construct) was replaced with the notion of there being distinct types of evidence that could be brought to bear on the validity of a given test,
depending on the intended purposes of the test. Broadly, these types of evidence
help establish that the test assesses as much as possible of what it should assess and
as little as possible of what it should not: in Messick’s language, this involves minimizing both construct underrepresentation and construct-irrelevant variance.
15
One of the more controversial elements of Messick’s perspective was the proposition that validation should explicitly involve a consideration of the consequences
of test interpretation and use. For example, if educational tests given to students are
used to help inform decisions about the retention and compensation of teachers,
claiming that the tests are valid for this purpose would involve demonstrating not
only that they measure the knowledge, skills, and abilities of students they claim to
measure, but also that using these tests as a basis for high-stakes decisions about
teachers has the intended positive consequences and does not have unforeseen negative consequences. This viewpoint could be taken as broadening the concept of
validity to include social and moral concerns in addition to more purely epistemic
concerns. Although Messick himself only proposed that the consequences of tests
could be used as indirect evidence of construct underrepresentation and constructirrelevant variance, other scholars such as Lorrie Shepard (1993) made stronger
proposals for the explicit consideration of consequences as a primary and independent source of validity evidence.
Messick’s view of validity has remained influential since its introduction, and is
arguably still the dominant conception of validity in the literature on educational
assessment and measurement. Using Messick’s definition of validity as a starting
point, scholars such as Michael Kane (1992) have argued that the activity of validation should consist of the construction and evaluation of an argument (or a set of
arguments) aimed at defending the appropriateness of a test for a particular, wellspecified use; the specification of such an argument then serves as an organizing
framework for the collection of forms of evidence necessary for its defense. Kane’s
argument-based approach is targeted to the practical problem of validation and not
a new theory about validity itself; this emphasis on validation rather than validity
reflects a shift in focus towards pragmatic, context-specific arguments tailored for
15 The term “influence properties” could be thought of as referring to sources of construct- (or
property-) irrelevant variance.
4 Philosophical perspectives on measurement
ity is a single property of a test. Second, Messick reframed validity from being a
claim about the true state of affairs (“a test measures what it claims to measure”) to
being a claim about the present state of available evidence, as judged by a particular
community—that is, from an ontological claim to an epistemic claim: “validity is an
integrated evaluative judgment of the degree to which empirical evidence and theoretical rationales support the adequacy and appropriateness of inferences and actions
based on test scores” (Messick, 1989: p. 13). The idea of distinct types of validity
(e.g., criterion, content, construct) was replaced with the notion of there being distinct types of evidence that could be brought to bear on the validity of a given test,
depending on the intended purposes of the test. Broadly, these types of evidence
help establish that the test assesses as much as possible of what it should assess and
as little as possible of what it should not: in Messick’s language, this involves minimizing both construct underrepresentation and construct-irrelevant variance.
15
One of the more controversial elements of Messick’s perspective was the proposition that validation should explicitly involve a consideration of the consequences
of test interpretation and use. For example, if educational tests given to students are
used to help inform decisions about the retention and compensation of teachers,
claiming that the tests are valid for this purpose would involve demonstrating not
only that they measure the knowledge, skills, and abilities of students they claim to
measure, but also that using these tests as a basis for high-stakes decisions about
teachers has the intended positive consequences and does not have unforeseen negative consequences. This viewpoint could be taken as broadening the concept of
validity to include social and moral concerns in addition to more purely epistemic
concerns. Although Messick himself only proposed that the consequences of tests
could be used as indirect evidence of construct underrepresentation and constructirrelevant variance, other scholars such as Lorrie Shepard (1993) made stronger
proposals for the explicit consideration of consequences as a primary and independent source of validity evidence.
Messick’s view of validity has remained influential since its introduction, and is
arguably still the dominant conception of validity in the literature on educational
assessment and measurement. Using Messick’s definition of validity as a starting
point, scholars such as Michael Kane (1992) have argued that the activity of validation should consist of the construction and evaluation of an argument (or a set of
arguments) aimed at defending the appropriateness of a test for a particular, wellspecified use; the specification of such an argument then serves as an organizing
framework for the collection of forms of evidence necessary for its defense. Kane’s
argument-based approach is targeted to the practical problem of validation and not
a new theory about validity itself; this emphasis on validation rather than validity
reflects a shift in focus towards pragmatic, context-specific arguments tailored for
15 The term “influence properties” could be thought of as referring to sources of construct- (or
property-) irrelevant variance.
4 Philosophical perspectives on measurement
