164
assess whether two given readers have indistinguishable RCAs (in analogy with the
comparison depicted in Fig. 6.2). For example, the two readers could be asked to
discuss the contents of a text passage with a human judge, and the judge could then
rate the reader’s relative RCAs. Now, unaided human judges may not have sufficient
resolution to discriminate RCA beyond rough ordinal classes (e.g., very little comprehension, text comprehension, literal comprehension, inferential comprehension),
so that one could, subject to the assumption that one used the same human judge,
consider that RCA is at most an ordinal property. Apart from concerns that this may
be assuming a weaker scale of RCA than possible, there are clearly serious issues of
subjectivity at play in this situation: Did the judge ask the same questions of the two
readers, did the judge rate the responses to the questions “fairly”, and would a different human judge be consistent with this one?
A key step forward was the implementation of standardized reading tests (Kelly,
1916; see Sects. 1.2.2 and 3.3.1), where readers would (1) read a text passage and
(2) answer a fixed set of questions (called in this context “items”
25
) about the contents of the passage; and then (3) their answers would be judged as correct or incorrect, and (4) readers would be given sum scores (e.g., the total number of items that
they answered correctly) on the reading comprehension test. Here readers who had
the same sum score would be indistinguishable with respect to their RCAs as measured by that test. Again, this would result in an ordinal scale (i.e., the readers who
scored 0, the readers who scored 1, …, the readers who scored K, for a test composed of K items), though, depending on the number of items in the set, there would
be a finer grain size than in the previous paragraph (i.e., as many levels as there are
different sum scores). This approach does address some of the subjectivity issues
raised by the previous approach: the same questions are asked of each reader, and,
with a suitable standardized mode of item response scoring, the variations due to
different human judges can be reduced, if not eliminated altogether. However, what
is not directly addressed are the issues of (a) the selection of text passages, and (b)
the selection of questions about those passages. Suppose, however, that one was
prepared to overlook these last two issues: one might convince oneself that the specific text passages and questions included in the test were acceptably suitable for all
the applications that were envisaged for the reading comprehension test. In that
case, one could adopt a norm-referenced approach to developing a scale (see Sect.
6.3.4), where the cumulative percentages of readers from a sample from a given
reference population (say, Grade 6 readers from X state in the year 20YZ) were used
to establish a mapping from the RCA scores on the test to percentiles of the sample.
This makes possible the so-called equipercentile equating to the (similarly calculated) results of other reading comprehension tests.
Thus, at this point in the account, the conclusion is then that values of RCA are
individual abilities identified as elements in an ordinal scale. It is interesting to note
that the sum scores which are the indexes used for the ranks can also be thought of
25 As we discuss in Chap. 7, each question of a test operates as a transducer, in this case transforming the RCA of a reader to a score.
6 Values, scales, and the existence of properties
assess whether two given readers have indistinguishable RCAs (in analogy with the
comparison depicted in Fig. 6.2). For example, the two readers could be asked to
discuss the contents of a text passage with a human judge, and the judge could then
rate the reader’s relative RCAs. Now, unaided human judges may not have sufficient
resolution to discriminate RCA beyond rough ordinal classes (e.g., very little comprehension, text comprehension, literal comprehension, inferential comprehension),
so that one could, subject to the assumption that one used the same human judge,
consider that RCA is at most an ordinal property. Apart from concerns that this may
be assuming a weaker scale of RCA than possible, there are clearly serious issues of
subjectivity at play in this situation: Did the judge ask the same questions of the two
readers, did the judge rate the responses to the questions “fairly”, and would a different human judge be consistent with this one?
A key step forward was the implementation of standardized reading tests (Kelly,
1916; see Sects. 1.2.2 and 3.3.1), where readers would (1) read a text passage and
(2) answer a fixed set of questions (called in this context “items”
25
) about the contents of the passage; and then (3) their answers would be judged as correct or incorrect, and (4) readers would be given sum scores (e.g., the total number of items that
they answered correctly) on the reading comprehension test. Here readers who had
the same sum score would be indistinguishable with respect to their RCAs as measured by that test. Again, this would result in an ordinal scale (i.e., the readers who
scored 0, the readers who scored 1, …, the readers who scored K, for a test composed of K items), though, depending on the number of items in the set, there would
be a finer grain size than in the previous paragraph (i.e., as many levels as there are
different sum scores). This approach does address some of the subjectivity issues
raised by the previous approach: the same questions are asked of each reader, and,
with a suitable standardized mode of item response scoring, the variations due to
different human judges can be reduced, if not eliminated altogether. However, what
is not directly addressed are the issues of (a) the selection of text passages, and (b)
the selection of questions about those passages. Suppose, however, that one was
prepared to overlook these last two issues: one might convince oneself that the specific text passages and questions included in the test were acceptably suitable for all
the applications that were envisaged for the reading comprehension test. In that
case, one could adopt a norm-referenced approach to developing a scale (see Sect.
6.3.4), where the cumulative percentages of readers from a sample from a given
reference population (say, Grade 6 readers from X state in the year 20YZ) were used
to establish a mapping from the RCA scores on the test to percentiles of the sample.
This makes possible the so-called equipercentile equating to the (similarly calculated) results of other reading comprehension tests.
Thus, at this point in the account, the conclusion is then that values of RCA are
individual abilities identified as elements in an ordinal scale. It is interesting to note
that the sum scores which are the indexes used for the ranks can also be thought of
25 As we discuss in Chap. 7, each question of a test operates as a transducer, in this case transforming the RCA of a reader to a score.
6 Values, scales, and the existence of properties
