9 Genomic Techniques and How to Apply Them to Marine Questions
343
are homologous. However, homology is not in all cases the reason for sequence
similarity. Short sequences may be similar by chance or sequences may be similar
because both were selected to bind to a particular protein, such as a transcription factor. Two principal types of homology can be distinguished. Orthologous sequences
are homologous sequences that were separated by a speciation event. A gene that
existed in an ancestral species, which then diverged into two species, will then exist
as two copies. These copies are called orthologues and will typically have the same
or a similar function.
Homologous sequences that have been separated by a gene duplication event
in an ancestral genome are called paralogous. Paralogous genes may mutate and
acquire new functions because the selective pressure is reduced. The three terms
and their relation are depicted in Fig. 9.4.
Fig. 9.4 Relation of the terms homology, orthology and paralogy. Genes that have a common
ancestor are homologous, orthologous genes are separated by a speciation event. A gene duplication event generates two paralogous copies of a single gene. This figure is based on the figure at
the following website: http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/Orthology.html
Functional annotation of genomes can be described as the process of identifying the functions of particular regions of sequence data that would otherwise be
almost devoid of information (Overbeek et al. 2005). By assigning functions to coding regions, the understanding of a complete genome is extended. The most direct
and reliable method to obtain the function of a coding region is to carry out biological experiments. This approach is time consuming and expensive and therefore not
practical for the vast amount of genomic data that has been generated.
Proteins that are encoded by genes with similar sequences often have a similar
three-dimensional structure. Since the three dimensional structure determines the
function of a protein, the assumption can be made that genes with a similar sequence
encode proteins with a similar function. The connection between the sequence of a
gene and the function of the corresponding gene product can be exploited to obtain a
functional annotation for an unknown gene. For this purpose, its nucleotide or amino
acid sequence is compared to the sequences of genes with known and verified function. If a relevant sequence similarity can be found, the function of the known gene
343
are homologous. However, homology is not in all cases the reason for sequence
similarity. Short sequences may be similar by chance or sequences may be similar
because both were selected to bind to a particular protein, such as a transcription factor. Two principal types of homology can be distinguished. Orthologous sequences
are homologous sequences that were separated by a speciation event. A gene that
existed in an ancestral species, which then diverged into two species, will then exist
as two copies. These copies are called orthologues and will typically have the same
or a similar function.
Homologous sequences that have been separated by a gene duplication event
in an ancestral genome are called paralogous. Paralogous genes may mutate and
acquire new functions because the selective pressure is reduced. The three terms
and their relation are depicted in Fig. 9.4.
Fig. 9.4 Relation of the terms homology, orthology and paralogy. Genes that have a common
ancestor are homologous, orthologous genes are separated by a speciation event. A gene duplication event generates two paralogous copies of a single gene. This figure is based on the figure at
the following website: http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/Orthology.html
Functional annotation of genomes can be described as the process of identifying the functions of particular regions of sequence data that would otherwise be
almost devoid of information (Overbeek et al. 2005). By assigning functions to coding regions, the understanding of a complete genome is extended. The most direct
and reliable method to obtain the function of a coding region is to carry out biological experiments. This approach is time consuming and expensive and therefore not
practical for the vast amount of genomic data that has been generated.
Proteins that are encoded by genes with similar sequences often have a similar
three-dimensional structure. Since the three dimensional structure determines the
function of a protein, the assumption can be made that genes with a similar sequence
encode proteins with a similar function. The connection between the sequence of a
gene and the function of the corresponding gene product can be exploited to obtain a
functional annotation for an unknown gene. For this purpose, its nucleotide or amino
acid sequence is compared to the sequences of genes with known and verified function. If a relevant sequence similarity can be found, the function of the known gene
