91
specific audiences and circumstances. The argument-based approach emphasizes
that any validation effort begins with a clear statement of the proposed uses and
interpretations of a test, and that if tests are used for purposes other than those
originally intended, this requires a reexamination of the argument or the development of an entirely new argument. According to this view, the consequences of
testing would play a central role in a validity argument for a given test insofar as the
proposed use of the test implies an intention for certain consequences to happen (or
not to happen) as a result.
Messick’s and Kane’s perspectives have been influential in shaping recent editions of the Standards for Educational and Psychological Testing (e.g., AERA,
2014; see also Wilson, 2005) which aim to provide guidance to practitioners on the
construction of convincing arguments for the adequacy and appropriateness of tests
for given purposes, and describe different types of evidence that might be brought
to bear on such arguments. Echoing Messick, the Standards define as
“the degree to which evidence and theory support the interpretations of test scores
for proposed uses of tests”. The Standards go on to specify that sources of such
evidence may come from several sources, related in particular to (a) the content of
the test (evaluated by, e.g., expert judgment of the alignment of test items and the
property, the content and formats of the items, and the materials supporting test
interpretations); (b) the cognitive processes engaged in by examinees when responding to test items (evaluated by, e.g., interviews, observation, or self-report); (c) the
extent to which empirical patterns of test results are consistent with theory- based
expectations; (d) the extent to which patterns of empirical relations between the test
results and other forms of quantified information are consistent with theory-based
expectations; and (e) patterns of real-world outcomes (“consequences”).
Of course, the degree of support for any proposition is logically independent
from the truth of that proposition, and thus the conception of validity could be
regarded as essentially legalistic rather than scientific.
16
Although the term “measurement” is used throughout the Standards, its meaning is never defined; it seems
to be used interchangeably with the terms “testing” and “assessment”. Thus, as the
concept of validity has expanded, the concept of measurement has arguably been
buried, and it is not always obvious how or even if they are intended to relate (for an
expanded discussion, see Maul, 2014).
It could be noted that considering tests or assessments as measuring instruments
is only one among many possible interpretations, and tests are routinely put to uses
that, strictly speaking, do not appear to require that any measurement take place at
all; for example, the function of a school exam might simply be to inspire students
16 A legalistic conception of validity “operationalizes the concept in a way that makes it clear for
test developers what the exact standard for validity is: they have to convince the jury. This bears all
the marks of a licensing procedure. However, for scientific research, licensing procedures do not
suffice. Truth cannot be […] equated to amounts of evidence” as noted by Borsboom (2012: p. 40).
4.3 The concept of validity in psychosocial measurement
specific audiences and circumstances. The argument-based approach emphasizes
that any validation effort begins with a clear statement of the proposed uses and
interpretations of a test, and that if tests are used for purposes other than those
originally intended, this requires a reexamination of the argument or the development of an entirely new argument. According to this view, the consequences of
testing would play a central role in a validity argument for a given test insofar as the
proposed use of the test implies an intention for certain consequences to happen (or
not to happen) as a result.
Messick’s and Kane’s perspectives have been influential in shaping recent editions of the Standards for Educational and Psychological Testing (e.g., AERA,
2014; see also Wilson, 2005) which aim to provide guidance to practitioners on the
construction of convincing arguments for the adequacy and appropriateness of tests
for given purposes, and describe different types of evidence that might be brought
to bear on such arguments. Echoing Messick, the Standards define
“the degree to which evidence and theory support the interpretations of test scores
for proposed uses of tests”. The Standards go on to specify that sources of such
evidence may come from several sources, related in particular to (a) the content of
the test (evaluated by, e.g., expert judgment of the alignment of test items and the
property, the content and formats of the items, and the materials supporting test
interpretations); (b) the cognitive processes engaged in by examinees when responding to test items (evaluated by, e.g., interviews, observation, or self-report); (c) the
extent to which empirical patterns of test results are consistent with theory- based
expectations; (d) the extent to which patterns of empirical relations between the test
results and other forms of quantified information are consistent with theory-based
expectations; and (e) patterns of real-world outcomes (“consequences”).
Of course, the degree of support for any proposition is logically independent
from the truth of that proposition, and thus the conception of validity could be
regarded as essentially legalistic rather than scientific.
16
Although the term “measurement” is used throughout the Standards, its meaning is never defined; it seems
to be used interchangeably with the terms “testing” and “assessment”. Thus, as the
concept of validity has expanded, the concept of measurement has arguably been
buried, and it is not always obvious how or even if they are intended to relate (for an
expanded discussion, see Maul, 2014).
It could be noted that considering tests or assessments as measuring instruments
is only one among many possible interpretations, and tests are routinely put to uses
that, strictly speaking, do not appear to require that any measurement take place at
all; for example, the function of a school exam might simply be to inspire students
16 A legalistic conception of validity “operationalizes the concept in a way that makes it clear for
test developers what the exact standard for validity is: they have to convince the jury. This bears all
the marks of a licensing procedure. However, for scientific research, licensing procedures do not
suffice. Truth cannot be […] equated to amounts of evidence” as noted by Borsboom (2012: p. 40).
4.3 The concept of validity in psychosocial measurement
