87
Standards for Educational and Psychological Testing, AERA, 2014: p. 11; for
recent comprehensive treatments of the topic of validity from philosophical perspectives, see Markus & Borsboom, 2013; Slaney, 2017). Moreover, as has been
argued by, e.g., Borsboom (2006, 2009), Michell (2009), Maul, Torres Irribarra, and
Wilson (2016), and Slaney (2017), the actual practice of psychosocial measurement
seems to be largely disconnected from and unconcerned with the philosophical conceptions of measurement described in this chapter; indeed, within the mainstream
literature on psychosocial measurement, one would be hard-pressed to find serious
engagement with even basic philosophical questions such as what measurement is,
even in authoritative sources such as the previously mentioned Standards (Maul,
2014; see also Borsboom, 2009; Michell, 1997). Conversely, the literature on validity has developed more directly in tandem with the practice of psychosocial measurement (Newton & Shaw, 2014), and thus is both more reactive to and influential
on such practices. Thus, to more thoroughly appreciate how thinking about measurement has developed in the human sciences—also in the service of our larger
goal of understanding measurement across the sciences—it will be useful to briefly
review the literature on validity and validation.
There have been several distinct phases in the history of thinking and discourse
about validity over the twentieth and twenty-first centuries, often dovetailing with
the history of thinking and discourse about measurement (as reviewed in the previous sections), and science and knowledge even more generally. Even today, despite
wide agreement regarding its importance, there is no single conception of validity
universally accepted in the scholarly and professional communities in the human
sciences, and there remains considerable controversy surrounding its definition, as
well as about many related concepts and terms. This next section briefly reviews this
history (see also Maul, 2018).
4.3.1 Early perspectives on validity
One of the earliest proposed conceptions of validity, and one that remains popular
among many contemporary scholars and practitioners, is that validity is about the
extent to which an instrument measures what it is claimed to measure.
11
Against the
backdrop of logical positivism, behaviorism, and operationalism (as discussed in
the previous sections), formal accounts of validity in the early twentieth century—
due to scholars such as Truman L. Kelley and Edward E. Cureton—operationalized
11 Oftentimes validity is introduced alongside the concept of reliability, which usually refers to the
extent to which measurement results are free from random sources of measurement error. Although
some sources (e.g., Moss, 1994) describe reliability and validity as separate, complementary
issues, most contemporary descriptions emphasize that reliability is a precondition for validity
rather than a separate issue. As was discussed in Sect. 3.2.1, this usage of the terms “reliability”
and “validity” then seems to map fairly closely onto what the VIM refers to as “precision” and
“accuracy”, respectively (JCGM, 2012: 2.15 and 2.13).
4.3 The concept of validity in psychosocial measurement
Standards for Educational and Psychological Testing, AERA, 2014: p. 11; for
recent comprehensive treatments of the topic of validity from philosophical perspectives, see Markus & Borsboom, 2013; Slaney, 2017). Moreover, as has been
argued by, e.g., Borsboom (2006, 2009), Michell (2009), Maul, Torres Irribarra, and
Wilson (2016), and Slaney (2017), the actual practice of psychosocial measurement
seems to be largely disconnected from and unconcerned with the philosophical conceptions of measurement described in this chapter; indeed, within the mainstream
literature on psychosocial measurement, one would be hard-pressed to find serious
engagement with even basic philosophical questions such as what measurement is,
even in authoritative sources such as the previously mentioned Standards (Maul,
2014; see also Borsboom, 2009; Michell, 1997). Conversely, the literature on validity has developed more directly in tandem with the practice of psychosocial measurement (Newton & Shaw, 2014), and thus is both more reactive to and influential
on such practices. Thus, to more thoroughly appreciate how thinking about measurement has developed in the human sciences—also in the service of our larger
goal of understanding measurement across the sciences—it will be useful to briefly
review the literature on validity and validation.
There have been several distinct phases in the history of thinking and discourse
about validity over the twentieth and twenty-first centuries, often dovetailing with
the history of thinking and discourse about measurement (as reviewed in the previous sections), and science and knowledge even more generally. Even today, despite
wide agreement regarding its importance, there is no single conception of validity
universally accepted in the scholarly and professional communities in the human
sciences, and there remains considerable controversy surrounding its definition, as
well as about many related concepts and terms. This next section briefly reviews this
history (see also Maul, 2018).
4.3.1 Early perspectives on validity
One of the earliest proposed conceptions of validity, and one that remains popular
among many contemporary scholars and practitioners, is that validity is about the
extent to which an instrument measures what it is claimed to measure.
11
Against the
backdrop of logical positivism, behaviorism, and operationalism (as discussed in
the previous sections), formal accounts of validity in the early twentieth century—
due to scholars such as Truman L. Kelley and Edward E. Cureton—operationalized
11 Oftentimes validity is introduced alongside the concept of reliability, which usually refers to the
extent to which measurement results are free from random sources of measurement error. Although
some sources (e.g., Moss, 1994) describe reliability and validity as separate, complementary
issues, most contemporary descriptions emphasize that reliability is a precondition for validity
rather than a separate issue. As was discussed in Sect. 3.2.1, this usage of the terms “reliability”
and “validity” then seems to map fairly closely onto what the VIM refers to as “precision” and
“accuracy”, respectively (JCGM, 2012: 2.15 and 2.13).
4.3 The concept of validity in psychosocial measurement
