6. Edit neospora.ini to match the following:
work_dir="install_dir/vacceed"
species_dir="neospora"
email_url="your_email@address" (user e-mail address)
proteome_fasta="proteome.fasta" (protein sequence file as per
step 1)
prot_id_prefix="xx" (needs to match the sequence identifier as
per step 1)
7. Modify the [Resources] in neospora.ini, if required. That is,
remove any resource names between VALIDATE and EVIDENCE that are not required (e.g., MHCI and MHCII).
8. Change directory to install_dir/vacceed/start in a commandline terminal.
9. Enter the command: perl startup nc (where “nc” is as per step
5).
10. Check results in “vaccine_candidates” in install_dir/vacceed/
neospora/proteome.
3.3 Creating
Pathogen Specific
Training Data
Training data here is essentially the collection of predicted evidence
(referred to henceforth as evidence profiles) from the seven bioinformatics programs for those proteins known to be positive or
negative. A training data file called “train_profiles” is provided
with the Vacceed package as part of the T. gondii sample data (see
Note 9). A previous study [8] tested Vacceed with different evidence profiles compiled from different eukaryotic species. It concluded that there is no fundamental difference in evidence profile
patterns; for example, a model trained on one species can be used to
classify proteins from another. This is because the bioinformatics
programs are designed or ML trained for eukaryotes in general.
Therefore, the creation of a pathogen specific training dataset is not
a mandatory step. However, an ideal training dataset is one that
contains the greatest variety of evidence profiles (see Note 10)
irrespective of the source species; for example, quality and variety
are indisputably the most important factors that impact the accuracy of ML algorithms [8]. A new or amended training file is
recommended under any of the following circumstances: a bioinformatics program is upgraded, that is, it has improved accuracy;
experimentally proved immunogenic proteins become available;
and a new prediction program is added (see Subheading 3.5).
1. Collect as many proteins as possible for the target species that
are known to induce an immune response in the relevant host.
The proteins will represent the “positives” for the training file
(see Note 11).
34
Stephen J. Goodswen et al.
work_dir="install_dir/vacceed"
species_dir="neospora"
email_url="your_email@address" (user e-mail address)
proteome_fasta="proteome.fasta" (protein sequence file as per
step 1)
prot_id_prefix="xx" (needs to match the sequence identifier as
per step 1)
7. Modify the [Resources] in neospora.ini, if required. That is,
remove any resource names between VALIDATE and EVIDENCE that are not required (e.g., MHCI and MHCII).
8. Change directory to install_dir/vacceed/start in a commandline terminal.
9. Enter the command: perl startup nc (where “nc” is as per step
5).
10. Check results in “vaccine_candidates” in install_dir/vacceed/
neospora/proteome.
3.3 Creating
Pathogen Specific
Training Data
Training data here is essentially the collection of predicted evidence
(referred to henceforth as evidence profiles) from the seven bioinformatics programs for those proteins known to be positive or
negative. A training data file called “train_profiles” is provided
with the Vacceed package as part of the T. gondii sample data (see
Note 9). A previous study [8] tested Vacceed with different evidence profiles compiled from different eukaryotic species. It concluded that there is no fundamental difference in evidence profile
patterns; for example, a model trained on one species can be used to
classify proteins from another. This is because the bioinformatics
programs are designed or ML trained for eukaryotes in general.
Therefore, the creation of a pathogen specific training dataset is not
a mandatory step. However, an ideal training dataset is one that
contains the greatest variety of evidence profiles (see Note 10)
irrespective of the source species; for example, quality and variety
are indisputably the most important factors that impact the accuracy of ML algorithms [8]. A new or amended training file is
recommended under any of the following circumstances: a bioinformatics program is upgraded, that is, it has improved accuracy;
experimentally proved immunogenic proteins become available;
and a new prediction program is added (see Subheading 3.5).
1. Collect as many proteins as possible for the target species that
are known to induce an immune response in the relevant host.
The proteins will represent the “positives” for the training file
(see Note 11).
34
Stephen J. Goodswen et al.
