6. If Vacceed fails with the test data then it will inevitably fail with
any other data. The expected reason for the failure is installation issues of one or more of the third-party programs (see
Note 4). The log file may give clues as to which third-party
program(s) is the culprit.
7. The ML algorithms used are listed in the configuration file
under the header [Evidence] and the key “algorithms.”
8. It is recommended that all known pathogen proteins are processed irrespective of protein name or expected function. This
allows for an unbiased approach.
9. This training file contains 475 positives of mainly T. gondii
proteins (nine are N. caninum). A small selection of these
proteins are known to induce an immune response, but most
are proteins predicted to be membrane-associated or secreted,
that is, proteins exposed to the immune system. There are
501 T. gondii proteins representing negatives, which were
defined by the protein’s predicted subcellular location, that is,
neither membrane-associated nor secreted.
10. Variety, in this instance from a ML perspective, is having a
generalized selection of proteins in the training file that are
representative of all conceivable types of positive and negative
proteins. For example, with a limited selection, a ML algorithm
may not generalize to evidence profiles not seen when it was
learning (i.e., poorly predicts when given new data).
11. Finding training proteins for most species is not a trivial task.
The expectation is that a thorough search of the literature will
be required. Even then, there may still be an inadequate number of examples to create a training file. A suggested compromise is to use positive proteins from a closely related organism
or proteins “expected” to induce or not induce an immune
response. For instance, use proteins known to be exposed to
the immune system (e.g., membrane associated or secreted
proteins) for positives and nonexposed proteins (e.g., proteins
normally located in the interior of the organism) as negatives.
12. A drawback for collecting negative examples is that a protein
cannot definitively be defined a negative unless it has been
explicitly tested in a laboratory.
13. The same proteins should never be used for training and evaluation. This would introduce biased results. Typically, the
proteins are randomly divided into two sets: one set containing
the majority of data (e.g., 80%) for training and the other set
(e.g., 20%) used to evaluate the trained model’s performance.
14. k-Fold cross-validation is a resampling statistical method used
to estimate the performance of ML models. The “k” refers to
the number of groups that a given data sample is to be split; for
40
Stephen J. Goodswen et al.
Précédent

- 55/595

Suivant