210
relation between the total biomass (g of O. vulgaris within
thanks) and the amount of O. vulgaris eDNA detected
(p-value = 0.01261) was found in aquarium experiments.
The species was also detected by PCR in 7 of the 8 water
samples taken at sea, and successfully quantified by qPCR in
5 samples. This preliminary study and innovative method
opens very promising perspectives for developing quick and
cheap tools for the assessment of O. vulgaris distribution and
abundance in the sea. The method could help in a close future
for quantifying unseen and elusive marine species, thus contributing to establish sustainable fisheries.
5.2.4 Single-Cell Transcriptomics: Gotta Catch
‘em all!
Sabrina N. Kalita
1*
, Uwe John
1
, Nancy Kuehne
1
1
Alfred Wegener Institute Helmholtz Centre for Polar and
Marine Research, Am Handelshafen 12, 27570 Bremerhaven
DE
*corresponding author: Sabrina.Kalita@awi.de
Keywords: Single-cell, Transcriptomics, Sequencing,
Data analysis
Eukaryotic genomes are difficult to sequence and moreover to assemble and annotate as they contain large intergenic
regions, introns and repetitive DNA chunks. This means, that
even when the complete sequence of the genome is known, it
is often difficult to spot particular genes in the assembly. One
approach to conquer this problem is to examine all the messenger RNA molecules transcribed from the genome as they
directly connect to a gene function. So called transcriptomic
sequencing makes it possible to profile organisms without
detours as their coding regions are straightforwardly accessible and hereby revealing gene expression dynamics. More
importantly, transcriptomic approaches lead to reference
databases consisting of coding regions for further modeling.
Unfortunately, some of the cellular properties are masked due
to bulk and population-averaged samples retrieved from pure
cultures. But in recent years, low-input RNA-sequencing
methods have been adapted to work with single cells. Thus,
making it possible to study single cell states, quantify intrapopulation heterogeneity potentially revealing cell subtypes
or gene expression dynamics that were previously masked in
bulk measurements. Even more, the possibility to work with
one single cell eliminates the need of cell cultivation and
enables the opportunity to process a cell directly from the
environment. Although now robust and economically practical, it is still a challenge to develop sensitive, precise, and
reliable protocols that lead to whole transcriptome sequencing from single cells. Here, we will report on the process of
establishing the whole work- flow regarding single-cell transcriptomics in the laboratory, leading to an analytical evaluation of methods and processes for single-cell sequencing.
This will include a feasibility study whether the method is
suitable for on board analysis on research vessels (RV
HEINCKE; Spitsbergen, NO; August ‘17), as well as data
analysis, like gene annotations after Illumina sequencing.
5.2.5 Defining Gene Cluster Families
from Globally- Distributed Seawater Samples
Using Community Detection Methods
Alaina Weinheimer
1*
, Jorge C. Navarro-Muñoz
2
, Frank
Oliver
Glöckner
1,3
,
Marnix
Medema
2
,
Antonio
Fernandez-Guerra
1,3,4
1
Microbial Genomics and Bioinformatics Research
Group, Max Planck Institute for Marine Microbiology,
Bremen, Germany
2
Bioinformatics Group, Wageningen University, 6708PB
Wageningen, Netherlands.
3
Jacobs University Bremen gGmbH, Bremen, Germany
4
Oxford e-Research Centre (OeRC), University of Oxford,
Oxford, UK
*corresponding author: alainarw94@gmail.com
Keywords: Natural products, Network analyses, Seawater
metagenomics
The pharmaceutical and agricultural industries often
rely on natural products synthesized by bacteria, plants,
and other organisms. Once isolated from nature, these
compounds or enzymes can be mass produced through cultivation. Alternatively, some industries genetically engineer organisms for the synthesis of such products.
However, the process and results of doing so are typically
laborious and/or unpredictable. Thus, finding useful
enzymes and metabolites in nature is still an effective
means of natural product discovery. Though many ecosystems have been investigated extensively for these compounds and enzymes, the ocean remains widely unexplored.
The biosynthetic potential of organisms was examined in
seawater samples collected from oceans around the globe
by the TARA Oceans expedition. Genomes were then
assembled from the metagenomic seawater samples. These
genomes were searched for the presence of biosynthetic
gene clusters (BGCs) using the program antiSMASH,
identifying 1384 BGCs within 93 of the metagenomeassembled genomes (MAGs). The aim of this study was to
investigate the relatedness of these BGCs and define families of closely related BGCs in marine samples. Each BGC
was compared to each other using the program BiGSCAPE, which generates a similarity-based network.
Within this network, communities of BGCs were identified
by employing various community detection algorithms,
such as HDBSCAN and Louvain. Based on several metrics, such as entropy, the most informative community
detection algorithm was affinity propagation. This algoAppendices
relation between the total biomass (g of O. vulgaris within
thanks) and the amount of O. vulgaris eDNA detected
(p-value = 0.01261) was found in aquarium experiments.
The species was also detected by PCR in 7 of the 8 water
samples taken at sea, and successfully quantified by qPCR in
5 samples. This preliminary study and innovative method
opens very promising perspectives for developing quick and
cheap tools for the assessment of O. vulgaris distribution and
abundance in the sea. The method could help in a close future
for quantifying unseen and elusive marine species, thus contributing to establish sustainable fisheries.
5.2.4 Single-Cell Transcriptomics: Gotta Catch
‘em all!
Sabrina N. Kalita
1*
, Uwe John
1
, Nancy Kuehne
1
1
Alfred Wegener Institute Helmholtz Centre for Polar and
Marine Research, Am Handelshafen 12, 27570 Bremerhaven
DE
*corresponding author: Sabrina.Kalita@awi.de
Keywords: Single-cell, Transcriptomics, Sequencing,
Data analysis
Eukaryotic genomes are difficult to sequence and moreover to assemble and annotate as they contain large intergenic
regions, introns and repetitive DNA chunks. This means, that
even when the complete sequence of the genome is known, it
is often difficult to spot particular genes in the assembly. One
approach to conquer this problem is to examine all the messenger RNA molecules transcribed from the genome as they
directly connect to a gene function. So called transcriptomic
sequencing makes it possible to profile organisms without
detours as their coding regions are straightforwardly accessible and hereby revealing gene expression dynamics. More
importantly, transcriptomic approaches lead to reference
databases consisting of coding regions for further modeling.
Unfortunately, some of the cellular properties are masked due
to bulk and population-averaged samples retrieved from pure
cultures. But in recent years, low-input RNA-sequencing
methods have been adapted to work with single cells. Thus,
making it possible to study single cell states, quantify intrapopulation heterogeneity potentially revealing cell subtypes
or gene expression dynamics that were previously masked in
bulk measurements. Even more, the possibility to work with
one single cell eliminates the need of cell cultivation and
enables the opportunity to process a cell directly from the
environment. Although now robust and economically practical, it is still a challenge to develop sensitive, precise, and
reliable protocols that lead to whole transcriptome sequencing from single cells. Here, we will report on the process of
establishing the whole work- flow regarding single-cell transcriptomics in the laboratory, leading to an analytical evaluation of methods and processes for single-cell sequencing.
This will include a feasibility study whether the method is
suitable for on board analysis on research vessels (RV
HEINCKE; Spitsbergen, NO; August ‘17), as well as data
analysis, like gene annotations after Illumina sequencing.
5.2.5 Defining Gene Cluster Families
from Globally- Distributed Seawater Samples
Using Community Detection Methods
Alaina Weinheimer
1*
, Jorge C. Navarro-Muñoz
2
, Frank
Oliver
Glöckner
1,3
,
Marnix
Medema
2
,
Antonio
Fernandez-Guerra
1,3,4
1
Microbial Genomics and Bioinformatics Research
Group, Max Planck Institute for Marine Microbiology,
Bremen, Germany
2
Bioinformatics Group, Wageningen University, 6708PB
Wageningen, Netherlands.
3
Jacobs University Bremen gGmbH, Bremen, Germany
4
Oxford e-Research Centre (OeRC), University of Oxford,
Oxford, UK
*corresponding author: alainarw94@gmail.com
Keywords: Natural products, Network analyses, Seawater
metagenomics
The pharmaceutical and agricultural industries often
rely on natural products synthesized by bacteria, plants,
and other organisms. Once isolated from nature, these
compounds or enzymes can be mass produced through cultivation. Alternatively, some industries genetically engineer organisms for the synthesis of such products.
However, the process and results of doing so are typically
laborious and/or unpredictable. Thus, finding useful
enzymes and metabolites in nature is still an effective
means of natural product discovery. Though many ecosystems have been investigated extensively for these compounds and enzymes, the ocean remains widely unexplored.
The biosynthetic potential of organisms was examined in
seawater samples collected from oceans around the globe
by the TARA Oceans expedition. Genomes were then
assembled from the metagenomic seawater samples. These
genomes were searched for the presence of biosynthetic
gene clusters (BGCs) using the program antiSMASH,
identifying 1384 BGCs within 93 of the metagenomeassembled genomes (MAGs). The aim of this study was to
investigate the relatedness of these BGCs and define families of closely related BGCs in marine samples. Each BGC
was compared to each other using the program BiGSCAPE, which generates a similarity-based network.
Within this network, communities of BGCs were identified
by employing various community detection algorithms,
such as HDBSCAN and Louvain. Based on several metrics, such as entropy, the most informative community
detection algorithm was affinity propagation. This algoAppendices
