51
instrument with the measurand, the variability of the sample means converges to
zero, showing that the total influence of the aforementioned “small causes” is progressively reduced. Reasonably, then, the mean of a sufficiently big sample estimates “the mean value that the best possible instrument would have generated”.
However, in order to make this concept of “best possible instrument” operational, a second condition must be fulfilled: an instrument is expected to maintain its
calibration over time, so that, together with the calibration information, its indication values are sufficient to obtain appropriate values of the measurand. If the calibration information were not updated, the measurement results produced by the
no-longer-calibrated instrument would be systematically biased; critically, this bias
would not be revealed by the repeated application of the instrument itself, as it is
unaffected by sample size. In other words, the convergence to a target distribution is
not sufficient to obtain the value that would be generated by “the best possible
instrument”.
This is a delicate point: the quality of the information produced by measurement
is hindered by causes traditionally treated as belonging to two distinct and independent kinds, called random and systematic, respectively. This distinction can be functionally characterized: the observed variability in repeated measurements is treated
as being due to random causes, whereas systematic causes generate a bias which
remains constant across repeated measurements, whose results are then affected in
the same way by such causes. The consequence is that the effects of some systematic causes could be unobservable, and while there may be methods for reducing the
effects of random causes, by applying what we have called an empirical strategy,
nothing analogous is generally known for systematic causes. Furthermore, the consideration that the effects of these two kinds of causes manifest themselves in statistical versus nonstatistical ways led the authors of the GUM to conclude that they
“were to be combined in their own way and were to be reported separately (or when
a single number was required, combined in some specified way)” (JCGM, 2008a,
2008b: E.1.3).
On this basis a conceptualization was developed that considers the value generated by “the best possible instrument” to be an intrinsic feature of the measurand,
traditionally called the true value of the measurand and defined as “the value which
characterizes a quantity perfectly defined, in the conditions which exist when that
quantity is considered” (ISO, 1984: 1.18). The designation “error approach” has this
origin: due to its experimental component, measurement is unavoidably affected by
errors, understood as the difference between the measured value and the true value.
Under the assumption of the unknowability of true values, but with the aim of maintaining the operational applicability of the framework, this sharp characterization
has been sometimes weakened by instead considering “conventional true values”
(of course, the very concept of conventional truth is questionable, to say the least)
(see, e.g., ISO, 1984: 3.10) or “reference values” (see, e.g., JCGM, 2012: 2.16). The
philosophical justification of the claim of the very existence of a true value of an
empirical property is controversial, and we do not discuss it here further.
3.2 The quality of measurement and its results
instrument with the measurand, the variability of the sample means converges to
zero, showing that the total influence of the aforementioned “small causes” is progressively reduced. Reasonably, then, the mean of a sufficiently big sample estimates “the mean value that the best possible instrument would have generated”.
However, in order to make this concept of “best possible instrument” operational, a second condition must be fulfilled: an instrument is expected to maintain its
calibration over time, so that, together with the calibration information, its indication values are sufficient to obtain appropriate values of the measurand. If the calibration information were not updated, the measurement results produced by the
no-longer-calibrated instrument would be systematically biased; critically, this bias
would not be revealed by the repeated application of the instrument itself, as it is
unaffected by sample size. In other words, the convergence to a target distribution is
not sufficient to obtain the value that would be generated by “the best possible
instrument”.
This is a delicate point: the quality of the information produced by measurement
is hindered by causes traditionally treated as belonging to two distinct and independent kinds, called random and systematic, respectively. This distinction can be functionally characterized: the observed variability in repeated measurements is treated
as being due to random causes, whereas systematic causes generate a bias which
remains constant across repeated measurements, whose results are then affected in
the same way by such causes. The consequence is that the effects of some systematic causes could be unobservable, and while there may be methods for reducing the
effects of random causes, by applying what we have called an empirical strategy,
nothing analogous is generally known for systematic causes. Furthermore, the consideration that the effects of these two kinds of causes manifest themselves in statistical versus nonstatistical ways led the authors of the GUM to conclude that they
“were to be combined in their own way and were to be reported separately (or when
a single number was required, combined in some specified way)” (JCGM, 2008a,
2008b: E.1.3).
On this basis a conceptualization was developed that considers the value generated by “the best possible instrument” to be an intrinsic feature of the measurand,
traditionally called the true value of the measurand and defined as “the value which
characterizes a quantity perfectly defined, in the conditions which exist when that
quantity is considered” (ISO, 1984: 1.18). The designation “error approach” has this
origin: due to its experimental component, measurement is unavoidably affected by
errors, understood as the difference between the measured value and the true value.
Under the assumption of the unknowability of true values, but with the aim of maintaining the operational applicability of the framework, this sharp characterization
has been sometimes weakened by instead considering “conventional true values”
(of course, the very concept of conventional truth is questionable, to say the least)
(see, e.g., ISO, 1984: 3.10) or “reference values” (see, e.g., JCGM, 2012: 2.16). The
philosophical justification of the claim of the very existence of a true value of an
empirical property is controversial, and we do not discuss it here further.
3.2 The quality of measurement and its results
