P nf t
ð Þ ¼ e
À k pfl þ k ifl
ð
Þt
ð4:1Þ
The mean time to failure here is
MTTF nf ¼
1
k pfl þ k ifl
ð4:2Þ
If we assume that transient faults happen k times more often than permanent
faults, we get
P 1 t
ð Þ ¼ e
À 1 þ k
ð
Þ k pfl
ð
Þt
ð4:3Þ
MTTF 1 ¼
1
1 þ k
ð
Þk pfl
ð4:4Þ
As already shown above, k can be as high as 10
3
–10
5 especially for space-borne
systems.
If a system or processor is capable to (a) identify transient faults and (b) recover
from them, additional hardware for both of these two features is required, which of
course changes the reliability of the system.
Thus, it follows that
P ft t
ð Þ ¼ e
À 1 þ d i þ d r
ð
Þ1 þ k
ð
Þ k pfl t
ð4:5Þ
MTTF ft ¼
1
1 þ d i þ d r
ð
Þ1 þ k
ð
Þk pfl
ð4:6Þ
where d i reflects the additional hardware required for fault identification and d r the
required hardware for recovery. It is clear that the smaller d i and d r are, the higher
the reliability that can be achieved for a processor or memory subsystem.
However, the purpose of the newly introduced hardware is the identification and
toleration of transient faults is not fully achievable, we have to assume that the
system can detect and recover from a part of the transient faults. True, the absolute
number of faults and consequently also the permanent-to-transient fault ratio
becomes lower.
The system cannot, however, detect and recover all faults, thus we introduce a as
an indication of how successful the recovery is. a reflects the reduction of k using
fault tolerance mechanisms and must therefore be in the range of (0–1):
P ft t
ð Þ ¼ e
À 1 þ d i þ d r
ð
Þ1 þ ak
ð
Þ k pfl t
ð4:7Þ
MTTF ft ¼
1
1 þ d i þ d r
ð
Þ1 þ ak
ð
Þk pfl
ð4:8Þ
40
4 Generalized Algorithm of Fault Tolerance (GAFT)
Précédent

- 55/315

Suivant