non-coding RNA (lncRNA) wasn’t discovered
until 1990 (Brannan et al. 1990). These spliced
and polyadenylated RNAs function in epigenetic
regulation, the generation and sequestration of
miRNAs, and various other functions. While
most of the studies have been run in animals,
thousands of lncRNAs have been annotated in
plant genomes, including IPS1, which sequesters
miR399 with a non-cleavable target bulge in
response to phosphate starvation across many
plant species (Franco-Zorrilla et al. 2007).
The small RNAs in plants include sno, si, and
miRNAs, with the snoRNAs evolutionarily
conserved back to Archaea. They are produced
from their own RNA precursors, or introns,
which are cleaved by endonucleases and trimmed
by exonucleases, until only the protein bound
60–250 bp snoRNA remains; they then guide the
protein complex’s methylation and pseudouridylation of rRNAs in the nucleolus. It is
even hypothesized that snoRNAs gave rise to
miRNAs based on their similarity in processing
including some overlap of enzymes, their similar
hairpin structure, and combination of function
(Scott and Ono 2011). There have been reports of
snoRNAs with miRNA-like characteristics, and
vice versa, and even small RNAs with complete
sno and miRNA function in animals, plants, and
yeast. In plants, both miRNAs and siRNAs are
cut to 22 and 21nt lengths by dicer-like proteins
1 and 4, respectively, and loaded onto Ago1 in
the RISC, with the main difference being that an
RNA hairpin is processed into a miRNA for
mRNA gene suppression, while a dsRNA is
diced into many siRNAs for pathogen gene
silencing.
When the Spirodela polyrhiza genome was
published in 2014, prediction programs were
able to detect miRNA precursors through
homologous sequences and RNA folding software (Wang et al. 2014a). In strain 7498, all
miRBase plant mature sequences were mapped
back to the genome, and flanking sequences
analyzed by RNAfold and miRCheck (Denman
1993; Jones-Rhoades and Bartel 2004). The
search predicted 413 miRNAs belonging to 93
families. This survey based on DNA sequencing
aimed to provide all possible miRNA genes, for
comparison to other plant genomes, with the
eventual aim of detecting their activity in later
RNA-seq experiments.
The earliest attempt at sequencing and analyzing S. polyrhiza miRNAs predated the published genome. This experiment, run at Peking
University Shenzhen Graduate School, was run
on strain LT5a, isolated from Lake Tai, using
three populations grown in SH media for 1, 3,
and 5 days under control conditions. Using
18-31nt sRNA on a HiSeq 2000 Illumina platform, they sequenced 24 million reads, 3.5 of
which matched conserved miRNAs in miRBase,
and 7.6 million that were not annotated in GenBank or Rfam. These 7.6 million reads were
analyzed by the MIREAP program and validated
by Mfold to identify 41 predicted novel miRNAs
(Zuker 2003). A summary of this and the other
small RNA-seq experiments is available in
Table 16.1.
In strain 9509, conserved and novel miRNAs
were identified through small RNA-sequencing
and an analysis of read count and distribution
(Michael et al. 2017). The study used 10 sRNA
libraries from a SOLiD5500 sequencer, aligned
to the genome allowing 1 mismatch, and then
annotated if the candidate has a stable hairpin
structure, sufficient miR reads, more than 1 miR*
read, and a 2 or 3 nt 3′ overhang (Table 16.1).
They identified conserved miRNAs by checking
for a strong BLAST homology to not only the
mature, but also hairpin structures in miRBase.
Next they used the program TargetFinder with a
cutoff score of 4 to identify the predicted targets
(Fahlgren and Carrington 2010). These transcription and structural requirements lead to the
prediction of 59 conserved miRNAs in 22 families, and 29 novel miRNAs, with 29 of the
conserved and 25 of the novel miRNAs being
predicted to regulate 991 mRNA targets.
Alongside the miRNA prediction, they were
able to predict trans-acting siRNAs (tasiRNAs),
from the sRNA library using previously established criteria (Howell et al. 2007; Johnson et al.
2009). Reads matching cDNA and the corresponding genomic regions had miRNA results
filtered out, and then, 50nt candidate transcripts
were required to have over 100 reads, with over
158
P. Fourounjian
until 1990 (Brannan et al. 1990). These spliced
and polyadenylated RNAs function in epigenetic
regulation, the generation and sequestration of
miRNAs, and various other functions. While
most of the studies have been run in animals,
thousands of lncRNAs have been annotated in
plant genomes, including IPS1, which sequesters
miR399 with a non-cleavable target bulge in
response to phosphate starvation across many
plant species (Franco-Zorrilla et al. 2007).
The small RNAs in plants include sno, si, and
miRNAs, with the snoRNAs evolutionarily
conserved back to Archaea. They are produced
from their own RNA precursors, or introns,
which are cleaved by endonucleases and trimmed
by exonucleases, until only the protein bound
60–250 bp snoRNA remains; they then guide the
protein complex’s methylation and pseudouridylation of rRNAs in the nucleolus. It is
even hypothesized that snoRNAs gave rise to
miRNAs based on their similarity in processing
including some overlap of enzymes, their similar
hairpin structure, and combination of function
(Scott and Ono 2011). There have been reports of
snoRNAs with miRNA-like characteristics, and
vice versa, and even small RNAs with complete
sno and miRNA function in animals, plants, and
yeast. In plants, both miRNAs and siRNAs are
cut to 22 and 21nt lengths by dicer-like proteins
1 and 4, respectively, and loaded onto Ago1 in
the RISC, with the main difference being that an
RNA hairpin is processed into a miRNA for
mRNA gene suppression, while a dsRNA is
diced into many siRNAs for pathogen gene
silencing.
When the Spirodela polyrhiza genome was
published in 2014, prediction programs were
able to detect miRNA precursors through
homologous sequences and RNA folding software (Wang et al. 2014a). In strain 7498, all
miRBase plant mature sequences were mapped
back to the genome, and flanking sequences
analyzed by RNAfold and miRCheck (Denman
1993; Jones-Rhoades and Bartel 2004). The
search predicted 413 miRNAs belonging to 93
families. This survey based on DNA sequencing
aimed to provide all possible miRNA genes, for
comparison to other plant genomes, with the
eventual aim of detecting their activity in later
RNA-seq experiments.
The earliest attempt at sequencing and analyzing S. polyrhiza miRNAs predated the published genome. This experiment, run at Peking
University Shenzhen Graduate School, was run
on strain LT5a, isolated from Lake Tai, using
three populations grown in SH media for 1, 3,
and 5 days under control conditions. Using
18-31nt sRNA on a HiSeq 2000 Illumina platform, they sequenced 24 million reads, 3.5 of
which matched conserved miRNAs in miRBase,
and 7.6 million that were not annotated in GenBank or Rfam. These 7.6 million reads were
analyzed by the MIREAP program and validated
by Mfold to identify 41 predicted novel miRNAs
(Zuker 2003). A summary of this and the other
small RNA-seq experiments is available in
Table 16.1.
In strain 9509, conserved and novel miRNAs
were identified through small RNA-sequencing
and an analysis of read count and distribution
(Michael et al. 2017). The study used 10 sRNA
libraries from a SOLiD5500 sequencer, aligned
to the genome allowing 1 mismatch, and then
annotated if the candidate has a stable hairpin
structure, sufficient miR reads, more than 1 miR*
read, and a 2 or 3 nt 3′ overhang (Table 16.1).
They identified conserved miRNAs by checking
for a strong BLAST homology to not only the
mature, but also hairpin structures in miRBase.
Next they used the program TargetFinder with a
cutoff score of 4 to identify the predicted targets
(Fahlgren and Carrington 2010). These transcription and structural requirements lead to the
prediction of 59 conserved miRNAs in 22 families, and 29 novel miRNAs, with 29 of the
conserved and 25 of the novel miRNAs being
predicted to regulate 991 mRNA targets.
Alongside the miRNA prediction, they were
able to predict trans-acting siRNAs (tasiRNAs),
from the sRNA library using previously established criteria (Howell et al. 2007; Johnson et al.
2009). Reads matching cDNA and the corresponding genomic regions had miRNA results
filtered out, and then, 50nt candidate transcripts
were required to have over 100 reads, with over
158
P. Fourounjian
