Column
#
Content
Values/Format
1
Chromosome name
chr
[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,
X,Y,M] or GRC accession
a
2
Annotation source
[ENSEMBL,HAVANA]
3
Feature type
[gene,transcript,exon,CDS,UTR,start_codon,stop_codon,
Selenocysteine]
4
Genomic start location
Integer-value (1-based)
5
Genomic end location
Integer-value
6
Score (not used)
.
7
Genomic strand
[+,À]
8
Genomic phase (for
CDS features)
[0,1,2,.]
9
Additional information
as key-value pairs
see https://www.gencodegenes.org/pages/data_format.
html
GTF/GFF files, as well as FASTA files can be downloaded from databases like
GENCODE (https://www.gencodegenes.org/), ENSEMBL (https://www.ensembl.org/
downloads.html), UCSC (http://hgdownload.cse.ucsc.edu/downloads.html), etc. You
need those file formats for generation of a genome index together with the corresponding
FASTA file of the genome as it is described earlier in this Chapter in Sect. 7.2.2.
A more detailed description about GFF/GTF file formats can be found on https://www.
ensembl.org/info/website/upload/gff.html and many other websites, which make these
available for download.
Review Question 3
Why do we have to download *.fa and *.gtf/*.gff of a certain genome of interest?
7.2.7 BED
The BED format provides a simpler way of representing the features in a molecule. Each
line represents a feature in a molecule and it has only three required fields: name (of
chromosome or scaffold), start, and end. The BED format uses 0-based coordinates for the
starts and 1-based for the ends. Headers are allowed. Those lines should be preceded by #
and they will be ignored.
The first three columns in a BED file are required, additional columns are optional [7, 8].
If you display the first lines of a BED file in the terminal, it looks like this:
88
M. Kappelmann-Fenzl
Précédent

- 97/225

Suivant