88
the idea of an instrument measuring what it is claimed to measure in terms of the
statistical correlation between test scores (and therefore more generally instrument
indications, to use the metrological terminology) and a “criterion variable”, that is,
an external criterion that was believed to be related in some way to the property that
the test is measuring. For example, the validity of a job placement test might be
defined in terms of the statistical correlation between its scores and some kind of
quantified information about job performance, or a short version of a test might be
evaluated in terms of the correlation of its scores with those of a longer or more
thorough battery of tests. Such test-criterion correlation coefficients were sometimes referred to as validity coefficients.
12
In other contexts, tests are developed from sets of content specifications: for
example, for many educational tests, a primary goal is to ensure adequate instructional attention to a domain covered in a course. In such contexts, prediction of a
specific external criterion could be regarded as less important than ensuring that the
content of the test was representatively sampled from the domain of interest; this, in
turn, is primarily established via documentation of the test construction procedures
and through expert review. This led to a distinction between criterion-related validity (also sometimes called predictive validity) and content validity, initially thought
of as each applying to different types of tests.
4.3.2 Construct validity
Starting in the 1950s, scholars such as Lee Cronbach and Paul Meehl began to
observe that while the concepts of criterion validity and content validity seemed to
be appropriate for many tests, some other kinds of tests appeared to require something else. In particular, psychological properties such as personality characteristics
(e.g., aggression, conscientiousness) and broadly defined cognitive abilities (e.g.,
general intelligence) seemed difficult to operationalize in terms of either a specific
domain of content coverage or relations with specific external criteria. This led
Cronbach and Meehl (1955) to introduce the concept of construct validity, which
was understood primarily in terms of how the property measured by a given test
related to a network of other properties, that is, a nomic network (which they called
a nomological network), which could in turn be estimated by examining, under
specified conditions, the statistical correlations between scores on the test and
scores on other tests (or other quantified information about the relevant properties);
one could then examine the extent to which these correlations were consistent with
predictions made based on the theory of what the test measured.
13
12 The influence of operationalism seems clear: while today one might still speak of a correlation
coefficient as a tool for the evaluation of validity, speaking of such a correlation as definitional of
validity blurs the distinction between what we know and how we know it. See also Borsboom and
Mellenbergh (2004).
13 Cronbach and Meehl’s conception of nomological networks drew from Carnap’s (1950) project
4 Philosophical perspectives on measurement
Précédent

- 122/319

Suivant