2. One position: # of reads overlapping with one position.
3. Whole genome: # of sequenced bases/size of genome.
The necessary sequencing coverage strongly depends on the performed sequencing method
(WGS, WES, RNA-Seq, ChIP-Seq), the reference genome size, and gene expression
patterns. Recommendations of sequencing coverage regarding the sequencing method
are listed in Table 7.1 (“Sequencing Coverage”. illumina.com. Illumina education.
Retrieved 2017-10-02.).
Review Question 1
What coverage of a human genome will one get with one lane of HiSeq3000 in pairedend Sequencing mode (2x75), if 300M clusters were bound (human genome size: 3.2
GB)?
7.2.2 FASTA
FASTA format is a text-based format for representing either nucleotide sequences or
peptide sequences, in which nucleotides or amino acids are represented using a singleletter code. The simplicity of FASTA format makes it easy to manipulate and parse using
text-processing tools and scripting languages like R, Python, Ruby, and Perl. A sequence in
FASTA format begins with a single-line description, followed by lines of sequence data.
The so-called defline starts with a “>” symbol and can thus be distinguished from the
sequence data. A more detailed description of the FASTA format and its purpose within
BLAST search can be found on the NCBI website https://blast.ncbi.nlm.nih.gov/Blast.cgi?
CMD¼Web&PAGE_TYPE¼BlastDocs&DOC_TYPE¼BlastHelp.
FASTA format example:
Table 7.1 Sequencing
coverage recommendations for
some common sequencing
methods
Sequencing method
Recommended coverage
Whole genome sequencing
~30Â to 50Â (for human)
Whole-exome sequencing
~100Â
RNA sequencing
~20–50 Mio. reads/sample
ChIP-Seq
~100Â
7 NGS Data
81
Précédent

- 90/225

Suivant