82
F. Cheng
5.2.3 Reconstruction of the Human Protein-Protein
Interactome
There are several experimental strategies for mapping protein-protein interactions
(PPIs), such as yeast two-hybrid assay (Y2H) that measures direct physical interactions in cells and affinity purification mass spectrometry that measure the composition of protein complexes. Specifically, we can reconstruct the human proteinprotein interactome network by assembling various publicly available PPI data: (i)
binary, physical PPIs tested by high-throughput Y2H systems from public available high-quality Y2H datasets [60, 61]; (ii) High-quality PPIs from the published
protein structure databases, such as Interactome3D [62], Instruct [63], and Interactome INSIDER [64]; (iii) kinase-substrate interactions by literature-derived lowthroughput and high-throughput experiments from KinomeNetworkX [65], Human
Protein Resource Database (HPRD) [66], PhosphoNetworks [67, 68], PhosphositePlus [69], DbPTM 3.0 [70], and Phospho. ELM [71]; (iv) signaling network by
literature-derived low-throughput experiments as annotated in SignaLink2.0 [72]; (v)
protein complexes data identified by a robust affinity purification mass spectrometry
methodology collected from BioPlex V2.0 [73]; and (vi) carefully literature-curated
PPIs identified by affinity purification followed by mass spectrometry (AP-MS) and
by literature-derived low-throughput experiments from BioGRID [74], PINA [75],
HPRD [66], MINT [76], IntAct [77], and InnateDB [78]. The detailed bioinformatics
resources for human protein-protein interactions are provided in Table 5.2.
5.2.4 Collection of Disease-Associated Genes/Proteins
In general, we can integrate disease-gene annotation data from multiple commonly
used bioinformatics resources currently available (Table 5.2).
OMIM, The OMIM database (Online Mendelian Inheritance in Man, http://www.
omim.org/) [79] is a comprehensive collection covering literature-curated human
disease genes with high-quality experimental validation evidence.
CTD, The Comparative Toxicogenomics Database (http://ctdbase.org/) [80] provides information about interactions between chemicals and gene products, and their
association with various diseases. Here, only manually curated gene-disease interactions from the literature were used.
ClinVar, ClinVar is a public archive of relationships among sequence variation and
various human phenotypes (https://www.ncbi.nlm.nih.gov/clinvar/) [81]. To improve
the data quality, only clinically significant relationships among variants and disease
traits annotated in ClinVar can be used.
GWAS Catalog, The NHGRI-EBI Catalog of published genome-wide association studies (GWAS, https://www.ebi.ac.uk/gwas/) [82] provides unbiased (singlenucleotide polymorphism) SNP-trait associations with genome-wide significance.
Usually, a SNP-trait with genome-wide significance (p < 5 × 10
−8 ) will be used.
F. Cheng
5.2.3 Reconstruction of the Human Protein-Protein
Interactome
There are several experimental strategies for mapping protein-protein interactions
(PPIs), such as yeast two-hybrid assay (Y2H) that measures direct physical interactions in cells and affinity purification mass spectrometry that measure the composition of protein complexes. Specifically, we can reconstruct the human proteinprotein interactome network by assembling various publicly available PPI data: (i)
binary, physical PPIs tested by high-throughput Y2H systems from public available high-quality Y2H datasets [60, 61]; (ii) High-quality PPIs from the published
protein structure databases, such as Interactome3D [62], Instruct [63], and Interactome INSIDER [64]; (iii) kinase-substrate interactions by literature-derived lowthroughput and high-throughput experiments from KinomeNetworkX [65], Human
Protein Resource Database (HPRD) [66], PhosphoNetworks [67, 68], PhosphositePlus [69], DbPTM 3.0 [70], and Phospho. ELM [71]; (iv) signaling network by
literature-derived low-throughput experiments as annotated in SignaLink2.0 [72]; (v)
protein complexes data identified by a robust affinity purification mass spectrometry
methodology collected from BioPlex V2.0 [73]; and (vi) carefully literature-curated
PPIs identified by affinity purification followed by mass spectrometry (AP-MS) and
by literature-derived low-throughput experiments from BioGRID [74], PINA [75],
HPRD [66], MINT [76], IntAct [77], and InnateDB [78]. The detailed bioinformatics
resources for human protein-protein interactions are provided in Table 5.2.
5.2.4 Collection of Disease-Associated Genes/Proteins
In general, we can integrate disease-gene annotation data from multiple commonly
used bioinformatics resources currently available (Table 5.2).
OMIM, The OMIM database (Online Mendelian Inheritance in Man, http://www.
omim.org/) [79] is a comprehensive collection covering literature-curated human
disease genes with high-quality experimental validation evidence.
CTD, The Comparative Toxicogenomics Database (http://ctdbase.org/) [80] provides information about interactions between chemicals and gene products, and their
association with various diseases. Here, only manually curated gene-disease interactions from the literature were used.
ClinVar, ClinVar is a public archive of relationships among sequence variation and
various human phenotypes (https://www.ncbi.nlm.nih.gov/clinvar/) [81]. To improve
the data quality, only clinically significant relationships among variants and disease
traits annotated in ClinVar can be used.
GWAS Catalog, The NHGRI-EBI Catalog of published genome-wide association studies (GWAS, https://www.ebi.ac.uk/gwas/) [82] provides unbiased (singlenucleotide polymorphism) SNP-trait associations with genome-wide significance.
Usually, a SNP-trait with genome-wide significance (p < 5 × 10
−8 ) will be used.
