most likely (referred to henceforth as positives) from those least
likely (referred to henceforth as negatives) to induce an immune
response.
Vacceed is the collective name for a configurable pipeline of
linked bioinformatics programs, Perl scripts, R functions, and
Linux shell scripts [5]. It was inspired by the principles of reverse
vaccinology [6], whereby antigen discovery starts in silico using the
pathogen genome rather than the traditional culture-based method
of cultivating and dissecting the pathogen itself. Vacceed has been
designed to facilitate an automated, high-throughput computational approach to predict vaccine candidates against eukaryotic
pathogens given protein sequences [7]. The pipeline uses various
standalone bioinformatics programs to predict various protein
characteristics. Vacceed is grounded on the underlying premise
that there is an expected difference between the set of characteristics defining positives to those of negatives. These differences are
typically not apparent to an observer and hence applying a rulebased approach to distinguish proteins is not feasible. Conversely,
machine learning (ML) has the capacity to detect obscure differences. Vacceed uses a set of ML algorithms trained on protein
characteristics of known positives and negatives to distinguish if a
yet to be classified protein is a positive or negative [8]. So far,
Vacceed has been used in studies to predict vaccine candidates for
Neospora caninum [9] and Cystoisospora suis [10].
This chapter provides step by step descriptions of how to
configure and operate Vacceed for a eukaryotic pathogen of the
user’s choice. A prerequisite for pathogen choice, nonetheless, is a
substantial representation of the pathogen’s proteome in the form
of quality protein sequences.
2 Vacceed Core Background Information
Vacceed can be downloaded from https://github.com/goodswen/
vacceed/releases. The download package includes a comprehensive
Vacceed User Guide and sample data. Note that Vacceed has been
designed for a Linux operating system and has only been tested on
Red Hat Enterprise Linux 7.5 but is expected to work on most
Linux distributions.
Each data processing stage in the Vacceed pipeline is an independent resource, which is built from a central Linux shell script
encapsulating all programs needed to perform specific but related
tasks. Typical tasks include predicting a particular protein characteristic. By default, Vacceed uses seven bioinformatics programs to
predict protein characteristics: SignalP 5.0 [11] (predicts presence
and location of signal peptide cleavage sites using deep neural networks); WoLF PSORT 0.2 [12] and TargetP 2.0 [2] (predict
subcellular localization); TMHMM 2.0 [4] (predicts
30
Stephen J. Goodswen et al.
Précédent

- 45/595

Suivant