8.1
Introduction
Depending on the samples sequenced (human, mouse, etc.) you need to generate a Genome
Index of your reference genome before you are able to align your sequencing reads.
Therefore, usually the comprehensive gene annotation on the primary assembly
(chromosomes and scaffolds) sequence regions (PRI; .gtf file) and the nucleotide sequence
(PRI, FASTA file) of the genome release of interest (e.g., GRCh38) are downloaded.
Genome sequence and annotation files can be downloaded from various freely accessible
databases as listed below:
• GENCODE: https://www.gencodegenes.org
• UCSC Genome Browser: https://hgdownload.soe.ucsc.edu/downloads.html
• Ensembl: https://www.ensembl.org/info/data/ftp/index.html
• NCBI RefSeq: https://www.ncbi.nlm.nih.gov/refseq/
Once the Genome sequence and annotation files have been downloaded, a Genome
Index should be created. Each Genome Index has to be created by the software tool you are
using for sequence alignment. In this chapter, we focus on the STAR and Bowtie2 alignment
tools.
8.2
Generate Genome Index via STAR
A key limitation with STAR [1] is its requirement for large memory space. STAR requires at
least 30 GB to align to the human or mouse genomes. In order to generate the Genome
Index with STAR; first, create a directory for the index (e.g., GenomeIndices/Star/
GRCh38_index). Then, copy the genome FASTA and Gene Transfer Format (GTF) files
into this directory.
Example:
106
M. Kappelmann-Fenzl
Précédent

- 115/225

Suivant