10 Challenges in Certifying Small-Scale (IoT) Hardware Random Number Generators
173
Researchers have commented on the ambiguity of SP800-22’s hypothesis and
statistical output, stating that more descriptive test output is required. Zhu et al.
propose a Q-value, computed using test statistics prior to their consolidation to pvalues [601]. The proposed statistic is more sensitive to total variation distance
and Kullback-Leiber divergence. This overcomes some of the issues caused by
correlations between the non-χ 2 level 2 tests of SP800-22 [601].
Dieharder implements many of the SP800-22 tests. As a result, it shares many of
the criticisms levelled at SP800-22 [206]. TestU01 is a more recent battery aimed at
allowing researchers to develop and evaluate their own RNGs, especially TRNGs.
There is little critical literature regarding this battery at present, so the independence
of tests in TestU01 is an open question at this time. Turan et al. comment on the
presence of some tests that they have found to be correlated being implemented in
the Crush batteries of TestU01 [560].
The diversity of a test methodology is related to, but separate from, the
independence of tests. Where independence is a measure of how related the results
from a set of tests may be, diversity is a measure of how many methods of evaluation
are used in the analysis of an RNG. A common observation is that the isolated use
of p-values is insufficient to fully characterize the randomness (or lack thereof) of a
sequence. Research by Hurley-Smith et al. explores TRNG and QRNG in detail to
identify flaws that were not detected by the most commonly used test batteries. In
these analyses, test correlation and diversity are key topics.
10.4.1 Randomness Testing Under Data Collection
Constraints: Analyzing the DESFire EV1
The first of these in-depth analyses was conducted over the Mifare DESFire EV1,
an RFID card produced by NXP [379]. The DESFire EV1 is used as a part of the
Transport for London (TfL) Oyster card scheme, as well as other loyalty and ewallet schemes throughout Europe. As a device that can store cash value, it requires
robust security to foster trust among vendors and users. The EV1 has achieved an
EAL4+ certification, based on its full security implementation.
Table 10.3 shows the Dieharder results for 3 DESFire EV1 cards. As mentioned
previously, data collection from the EV1 is challenging, requiring 12 days to obtain
64 MB of data. As a result, this was the largest amount of data able to be collected. A
total of 100 cards were tested, with all 100 passing. The 3 cards shown in this table
show the p-values reported by the Dieharder tests for all tests that can be performed
on 64 MB of data without rewinds.
Card 3 shows a single failure of the Dieharder battery, for the count the ones test.
However, this was not reproduced by any other card that was tested. Therefore, it is
reasonable to conclude that the Dieharder battery does not identify any significant
degree of non-randomness in the tested sequences.
Table 10.4 shows the pass rates for NIST tests. All SP800-22 tests were used
over the EV1 samples we collected.
173
Researchers have commented on the ambiguity of SP800-22’s hypothesis and
statistical output, stating that more descriptive test output is required. Zhu et al.
propose a Q-value, computed using test statistics prior to their consolidation to pvalues [601]. The proposed statistic is more sensitive to total variation distance
and Kullback-Leiber divergence. This overcomes some of the issues caused by
correlations between the non-χ 2 level 2 tests of SP800-22 [601].
Dieharder implements many of the SP800-22 tests. As a result, it shares many of
the criticisms levelled at SP800-22 [206]. TestU01 is a more recent battery aimed at
allowing researchers to develop and evaluate their own RNGs, especially TRNGs.
There is little critical literature regarding this battery at present, so the independence
of tests in TestU01 is an open question at this time. Turan et al. comment on the
presence of some tests that they have found to be correlated being implemented in
the Crush batteries of TestU01 [560].
The diversity of a test methodology is related to, but separate from, the
independence of tests. Where independence is a measure of how related the results
from a set of tests may be, diversity is a measure of how many methods of evaluation
are used in the analysis of an RNG. A common observation is that the isolated use
of p-values is insufficient to fully characterize the randomness (or lack thereof) of a
sequence. Research by Hurley-Smith et al. explores TRNG and QRNG in detail to
identify flaws that were not detected by the most commonly used test batteries. In
these analyses, test correlation and diversity are key topics.
10.4.1 Randomness Testing Under Data Collection
Constraints: Analyzing the DESFire EV1
The first of these in-depth analyses was conducted over the Mifare DESFire EV1,
an RFID card produced by NXP [379]. The DESFire EV1 is used as a part of the
Transport for London (TfL) Oyster card scheme, as well as other loyalty and ewallet schemes throughout Europe. As a device that can store cash value, it requires
robust security to foster trust among vendors and users. The EV1 has achieved an
EAL4+ certification, based on its full security implementation.
Table 10.3 shows the Dieharder results for 3 DESFire EV1 cards. As mentioned
previously, data collection from the EV1 is challenging, requiring 12 days to obtain
64 MB of data. As a result, this was the largest amount of data able to be collected. A
total of 100 cards were tested, with all 100 passing. The 3 cards shown in this table
show the p-values reported by the Dieharder tests for all tests that can be performed
on 64 MB of data without rewinds.
Card 3 shows a single failure of the Dieharder battery, for the count the ones test.
However, this was not reproduced by any other card that was tested. Therefore, it is
reasonable to conclude that the Dieharder battery does not identify any significant
degree of non-randomness in the tested sequences.
Table 10.4 shows the pass rates for NIST tests. All SP800-22 tests were used
over the EV1 samples we collected.
