75
19. Primers for sequencing (see Subheading 3.8).
20. 1% agarose gel made with SB buffer (46 g/L boric acid, 8 g/L
sodium hydroxide), stained with, e.g., GelRed
™
(Biotium).
21. DNA ladder (e.g., GeneRuler 1 kb DNA ladder, Thermo
Scientific).
22. Agarose gel electrophoresis apparatus.
3 Methods
Given a reference SBP with the desired binding specificity, reconstruction of an ancestral SBP with improved thermostability and
the same binding specificity requires a dataset of protein sequences
that are homologous to the reference SBP (see Note 2). These
sequences can be identified using a BLAST (Basic Local Alignment
Search Tool) search of protein sequence databases, such as the
NCBI database of nonredundant protein sequences or the
UniProt-KB database. The protein sequences should be selected to
maximize sequence diversity and phylogenetic diversity, but should
have the same binding specificity as the reference SBP (see Notes 3
and 4). The sequences should be evenly distributed throughout
sequence space; in other words, clusters of sequences with very
high identity (>90%) and outlier sequences with low identity to
any other sequence in the dataset should be avoided. A small set of
outgroup sequences should also be selected for rooting the phylogenetic tree (see Note 5). In the case of SBPs, a suitable outgroup
would be an SBP homologous to the reference SBP, but with a
different binding specificity; for example, in our recent work,
anionic amino acid-binding proteins served as an outgroup for cationic and neutral amino acid-binding proteins [6].
As a guideline, between 50 and 250 sequences should be
selected for the phylogenetic analysis. If too many sequences are
included, the phylogenetic analysis will be computationally intensive, while if too few sequences are included, there may be insufficient data to reconstruct the ancestral sequences accurately, or the
sequence diversity within the dataset may be insufficient to make
the ancestral proteins more thermostable than the reference SBP.
Before the sequences can be used for phylogenetic analysis, they
must be collated in a multiple sequence alignment. In the phylogenetic analysis and subsequent reconstruction of ancestral sequences,
it is assumed that each column in the multiple sequence alignment
corresponds to a set of residues that originated from a common
ancestor. If the alignment is incorrect, this assumption is violated;
hence, the quality of the multiple sequence alignment is a critical
factor in the success of the project. The number of gaps should be
minimized by careful editing of the alignment to remove poorly
3.1 Sequence
Collection
3.2 Multiple
Sequence Alignment
Improving FRET Sensors by Ancestral Gene Resurrection
Précédent

- 80/332

Suivant