9 Genomic Techniques and How to Apply Them to Marine Questions
341
shows the location of structural domains on a sequence via the residue-by-residue
mapping between the PDB (Protein Data Bank, see also Section 9.3.5.4) chain(s)
and the UniProt sequences obtained using data from the MSD. Only the PDB chains
representing non-overlapping regions are shown. Although protein structure is more
difficult to determine than sequence, representative structures are currently known
for about 2,000 protein families. Structures are more conserved than sequences,
and often reveal evolutionary relationships, which are hidden at the sequence level.
Structural data is also essential for providing detailed insights into a protein’s
function, catalytic mechanism and interactions with other proteins.
How can these annotations of known proteins (within the InterPro database)
be used to investigate the functions of novel protein sequences? That is where
InterProScan comes into use. InterProScan (Quevillon et al. 2005) is a sequence
search tool, which searches against the signature databases (as opposed to BLAST
or FASTA which look for similarity to individual sequences in the databases;
see Section 9.3.3.1). The input protein sequence format should be either free
text/raw, in FASTA format or in the UniProt format. Input nucleotide sequence format used is either free text/raw, FASTA, or DDBJ/EMBL/GenBank format. You
can use InterProScan via the European Bioinformatics Institute (EBI) web interface (http://www.ebi.ac.uk/InterProScan/) or the standalone version of InterProScan
(ftp://ftp.ebi.ac.uk/pub/databases/interpro/iprscan/) by downloading and installing it
locally on your computer.
Two tutorials from the 2can website (the bioinformatics educational resource
from the EBI) provide help with using InterProScan as a tool to characterize and
annotate sequences: http://www.ebi.ac.uk/2can/tutorials/
(b) TMHMM and SignalP
TMHMM is dedicated to the identification of transmembrane proteins, whereas
SignalP detects signal peptides. Transmembrane TM proteins are polypeptide
chain(s) that pass through the lipid bilayer. They are involved in a wide variety
of cellular functions such as transport and inter and intra-cellular communication.
The majority of proteomes are predicted containing between 20 and 25% of transmembrane proteins. The two major classes are: alpha-helix bundle proteins, which
are found in any type of membrane, and beta-barrel proteins, which are present
mostly in the outer membranes (of gram-negative bacteria, but also mitochondria
or chloroplasts). The identification of TM domains is not easy as few 3D structures are available. There are currently less than 1,000 TM protein 3D structures
among the almost 52,000 3D structures available in the Protein data bank (PDB) as
of July 2008. Hydrophobic regions of precursor proteins are also often mistaken for
TM regions and aliphatic helices are not always recognized as a TM region. The
most reliable prediction program used for the detection of membrane-spanning segments and topology is TMHMM 2.0 (http://www.cbs.dtu.dk/services/TMHMM/),
which is based on a hidden Markov model. It predicts the location and orientation of alpha helical regions (Sonnhammer et al. 1998, Krogh et al. 2001).
Unfortunately, transmembrane topology prediction programs sometimes predict signal peptides as transmembrane segments near the N-terminus. A signal peptide
Précédent

- 352/410

Suivant