using SPRI beads, which have the added benefit of concentrating
the PCR product.
1.4 Data Acquisition
and Processing
After SPRI cleanup, the combined PCR products are ready for NGS
sequencing. We have used the Illumina MiSeq and HiSeq 2500
platforms, but with appropriately designed primers, others could
be used.
The output of most NGS platforms is one or a series of FASTQ
files, which contain in plain text a nucleotide sequence (“read”) and
associated quality scores for each molecular cluster. While other
applications, such as whole genome sequencing, require alignment
to reference genomes to identify sequence coordinates, the reads
from PROSPECT will contain the plate quadrant, well, and strain
barcodes at fixed positions within the reported sequence. Several
software packages exist to count barcodes from FASTQ files, which
can then be mapped back to strains, well positions, and compounds, yielding a table of counts with one row per combination
of plate, well, and strain and relevant metadata.
While barcode count data are a reproducible proxy for census of
a particular strain in a well [5], they, as counts, are subject to
intrinsic Poisson shot noise [12] as well as, like other highthroughput experiments, experimental batch effects [13]. Concensus GLM uses a Negative Binomial Generalized Linear Model
(GLM) to account for the statistical noise as well “modeling out”
batch effects using experimental metadata as covariates. This yields
a table reporting the log (fold-change) (LFC) in strain abundance
relative to a negative control, and its P-value, per combination of
strain and compound. The vector of LFCs across all the strains for a
given compound is known as the CGIP.
1.5 Data
Interpretation
It is informative to consider the relationship of LFC to strain fitness
(i.e., inverse doubling time). Since fitness is an intrinsic property of
the interaction between a hypomorph’s genetic disruption and
chemical treatment, it is independent of assay duration and strain
pool composition.
LFC can be related to strain fitness in several ways. Assuming
exponential growth throughout the assay duration t and that a
strong negative control such as rifampin instantaneously arrests
growth such that its LFC Rifampin reflects the inoculum composition,
the fitness w c of a particular strain when treated by compound c is
given by
w c ¼ LFC c À LFC Rifampin
À
Á =t
We can use LFC to interpret fitness while avoiding the assumption of exponential growth throughout the assay by instead looking
at the fitness w s,c of a given strain s in a particular compound
326
Eachan O. Johnson and Deborah T. Hung
the PCR product.
1.4 Data Acquisition
and Processing
After SPRI cleanup, the combined PCR products are ready for NGS
sequencing. We have used the Illumina MiSeq and HiSeq 2500
platforms, but with appropriately designed primers, others could
be used.
The output of most NGS platforms is one or a series of FASTQ
files, which contain in plain text a nucleotide sequence (“read”) and
associated quality scores for each molecular cluster. While other
applications, such as whole genome sequencing, require alignment
to reference genomes to identify sequence coordinates, the reads
from PROSPECT will contain the plate quadrant, well, and strain
barcodes at fixed positions within the reported sequence. Several
software packages exist to count barcodes from FASTQ files, which
can then be mapped back to strains, well positions, and compounds, yielding a table of counts with one row per combination
of plate, well, and strain and relevant metadata.
While barcode count data are a reproducible proxy for census of
a particular strain in a well [5], they, as counts, are subject to
intrinsic Poisson shot noise [12] as well as, like other highthroughput experiments, experimental batch effects [13]. Concensus GLM uses a Negative Binomial Generalized Linear Model
(GLM) to account for the statistical noise as well “modeling out”
batch effects using experimental metadata as covariates. This yields
a table reporting the log (fold-change) (LFC) in strain abundance
relative to a negative control, and its P-value, per combination of
strain and compound. The vector of LFCs across all the strains for a
given compound is known as the CGIP.
1.5 Data
Interpretation
It is informative to consider the relationship of LFC to strain fitness
(i.e., inverse doubling time). Since fitness is an intrinsic property of
the interaction between a hypomorph’s genetic disruption and
chemical treatment, it is independent of assay duration and strain
pool composition.
LFC can be related to strain fitness in several ways. Assuming
exponential growth throughout the assay duration t and that a
strong negative control such as rifampin instantaneously arrests
growth such that its LFC Rifampin reflects the inoculum composition,
the fitness w c of a particular strain when treated by compound c is
given by
w c ¼ LFC c À LFC Rifampin
À
Á =t
We can use LFC to interpret fitness while avoiding the assumption of exponential growth throughout the assay by instead looking
at the fitness w s,c of a given strain s in a particular compound
326
Eachan O. Johnson and Deborah T. Hung
