6000 and 150 bp PE run, roughly, 20 k cells can be sequenced at
50 k reads/cell to generate a single cell atlas capturing EMT.
Indeed, the multitude of scRNA-seq techniques and methods are
rapidly evolving, and as cost of scRNA-seq decreases, previous
technologies will surely become obsolete. Research groups continue to push the limits and cost efficiency of scRNA-seq with
methods like cell hashing that allow for “super loading” of cells,
and it will only drive the cost down.
5 Bioinformatic Analysis: Overview of Procedure for Analysis of Results
Once single cell libraries are prepared and samples have been
sequenced, the first step in analyzing the data is to create an
expression matrix from the raw sequencing output. First, your bcl
file should be demutiplexed using bcl2fastq to produce fastq files
that can be checked for read quality control. A pipeline should be
established early on, to identify what type of analyses will be performed (see Fig. 1 for a general ScRNA-seq pipeline that can be
adapted). Following sequencing, Unique Molecular Identifiers
(UMIs) should be extracted and reads demultiplexed using
UMI-tools or zUMIs [28, 29]. To perform a quality check on
reads, a common tool used is FastQC [30]. Once reads have been
checked for quality control, they should be trimmed if a sample has
poor base per sequence quality scores below 20, or if any exogenous
nucleotides such as adapters were introduced. A few commonly
used trimming tools are Trimmomatic, TrimGalore, and Cutadapt
[31–33]. Trimmed reads can then be mapped back to a reference
genome or transcriptome using a bulk RNA-seq aligner/pseudoaligner such as STAR/Kallisto or an aligner appropriate for your
research question [34, 35]. Once reads have been mapped to genes,
they are counted on a per gene and per cell basis to generate a single
cell gene expression matrix [28, 36]. This matrix has a row for each
cell and a column for each gene. The i, j entry encodes the number
of molecules of mRNA for gene j in cell i. Therefore, each row
encodes the expression profile of a cell as a point in a highdimensional gene expression space, where there is a dimension for
each gene.
With the expression matrix in hand, we are now ready to begin
visualizing, exploring and analyzing the data. We begin by visualizing the high-dimensional single cell gene expression profiles in two
or three dimensions. Some popular tools for visualizing single cell
datasets include force layout embedding (FLE), UMAP, and tSNE
[14, 15]. Instead of applying these tools directly to the single cell
expression data, it can be helpful to first reduce the dimensionality
from 20,000 down to ~1000 by selecting variable genes, and then
down to ~100 using principal components analysis (PCA) or diffusion maps. This gradual decrease in dimensionality can help extract
Methods for in Vivo EMT at Single Cell Resolution
309
50 k reads/cell to generate a single cell atlas capturing EMT.
Indeed, the multitude of scRNA-seq techniques and methods are
rapidly evolving, and as cost of scRNA-seq decreases, previous
technologies will surely become obsolete. Research groups continue to push the limits and cost efficiency of scRNA-seq with
methods like cell hashing that allow for “super loading” of cells,
and it will only drive the cost down.
5 Bioinformatic Analysis: Overview of Procedure for Analysis of Results
Once single cell libraries are prepared and samples have been
sequenced, the first step in analyzing the data is to create an
expression matrix from the raw sequencing output. First, your bcl
file should be demutiplexed using bcl2fastq to produce fastq files
that can be checked for read quality control. A pipeline should be
established early on, to identify what type of analyses will be performed (see Fig. 1 for a general ScRNA-seq pipeline that can be
adapted). Following sequencing, Unique Molecular Identifiers
(UMIs) should be extracted and reads demultiplexed using
UMI-tools or zUMIs [28, 29]. To perform a quality check on
reads, a common tool used is FastQC [30]. Once reads have been
checked for quality control, they should be trimmed if a sample has
poor base per sequence quality scores below 20, or if any exogenous
nucleotides such as adapters were introduced. A few commonly
used trimming tools are Trimmomatic, TrimGalore, and Cutadapt
[31–33]. Trimmed reads can then be mapped back to a reference
genome or transcriptome using a bulk RNA-seq aligner/pseudoaligner such as STAR/Kallisto or an aligner appropriate for your
research question [34, 35]. Once reads have been mapped to genes,
they are counted on a per gene and per cell basis to generate a single
cell gene expression matrix [28, 36]. This matrix has a row for each
cell and a column for each gene. The i, j entry encodes the number
of molecules of mRNA for gene j in cell i. Therefore, each row
encodes the expression profile of a cell as a point in a highdimensional gene expression space, where there is a dimension for
each gene.
With the expression matrix in hand, we are now ready to begin
visualizing, exploring and analyzing the data. We begin by visualizing the high-dimensional single cell gene expression profiles in two
or three dimensions. Some popular tools for visualizing single cell
datasets include force layout embedding (FLE), UMAP, and tSNE
[14, 15]. Instead of applying these tools directly to the single cell
expression data, it can be helpful to first reduce the dimensionality
from 20,000 down to ~1000 by selecting variable genes, and then
down to ~100 using principal components analysis (PCA) or diffusion maps. This gradual decrease in dimensionality can help extract
Methods for in Vivo EMT at Single Cell Resolution
309
