chromosomes. In addition, you can evaluate whether the pattern matches the experiment by
looking for specific different patterns:
• TFs: enrichment of reads near the TSS and distal regulatory elements.
• H3K4me3—enrichment of reads near TSS.
• H3K4me1/2, H3/H4ac, DNase—enrichment of reads near TSS and distal regulatory
elements.
• H3K36me3—enrichment of reads across gene bodies.
• H3K27me3—enrichment of reads near CpG Islands of inactive genes.
• H3K9me3—enrichment of reads across broad domains and repeat elements.
12.4 Copy Number Variation (CNV) of Input Samples
ChIP-Seq experiments generally show an enrichment of reads in regulatory regions of the
genome. Thus, it is important to also sequence a non-chipped genomic DNA of the same
sample (Input) to control if the ChIP was successful. The Input sample can be further used to
analyze and visualize possible copy number variation (CNV) of the Input sample. This can be
important in terms of deletions or duplications of whole chromosomes or parts of a chromosome, which would lead to a false discovery of enriched reads in regions with alterations in
copy numbers in the further analysis workflow. Consequently, no or low peaks (enrichment
of reads) should be detected in the Input samples compared to the chipped once.
Control-FREEC [7] is a tool we have already installed via
conda for detection of
CNV and allelic imbalances (LOH) in NGS data (Chap. 7). It automatically computes,
normalizes, segments copy number and beta allele frequency (BAF) profiles. Then it calls
CNV and LOH. For whole genome sequencing data analysis, like our Input sample, the
program can also use mappability data (files created by GEM (https://sourceforge.net/
projects/gemlibrary/files/gem-library/)) [7]. To be able to run Control-FreeC you have to
create a FreeC directory in your GenomeIndices folder with all .fa files of the reference
genome as well as a file with chromosome sizes. Moreover, the mappability file should also
be stored here (or elsewhere, but you should remember where). A detailed description of
the usage of Control-FreeC can be found on the following website: http://boevalab.inf.ethz.
ch/FREEC/tutorial.html. Moreover, the calculation of significance of Control-FreeC
predictions and the visualization of Control-FreeC´s output using R are described. The
required scripts are stored on GitHub (https://github.com/BoevaLab/FREEC).
12.5 Peak/Region Calling
In terms of ChIP-Seq reads, finding peaks/regions is one of the central goals and the same
basic principles apply as for other types of sequencing. In terms of a transcription factor
ChIP-Seq experiment one speaks of identifying “peaks,” in terms of histone modifications
184
M. Kappelmann-Fenzl
Précédent

- 191/225

Suivant