5.1
Introduction
Many NGS data analysis programs are written in Java, Perl, C++ or Python. Simply
running these programs does not require programming knowledge, however, it is indeed
helpful. Within the framework of this textbook bioinformatic details are not covered.
Anyway, to be able to perform NGS data analysis familiarity with R is important. For
those of you who are not familiar with R, it is highly recommended to take advantage of
some freely available web resources (e.g. https://www.bioconductor.org/help/coursematerials/2012/SeattleMay2012/).
5.2
Computer Setup for NGS Data Analysis
Next-generation sequencing analysis is a computationally demanding process. Your average laptop is probably not up to the challenge. Any Linux/Unix-based operating system
will work well, with large servers in mind. Typically, analysis algorithms will be
distributed by researchers in one of the three ways:
• Standalone program to run in a Linux/Unix environment (most common).
• Webserver where you upload data to be analyzed (i.e. Galaxy).
• R software package/library for the R computing environment.
In theory, any computer with enough RAM, hard drive space, and CPU power can be
used for analysis. In general, you will need:
• 16 Gb of RAM Minimum (better to have 96+ Gb).
• 500 Gb of disk space (better to have 10+ Tb of space).
• Fast CPU (better to have at least 8 cores, more the better).
• External storage.
Depending on the type of analysis you want to perform, the required RAM, hard drive
space, and CPU power may vary and for some analyses a conventional laptop is fully
sufficient. If you own a computer or lab server with the aforementioned properties, you can
start over setting up your computer environment. The command line analysis tools that we
demonstrate in this book run on the Linux/Unix operating system. For practicing NGS data
analysis, we will use several software tools in various formats:
• Binary executable code that can be run directly on the target computer.
• Source code that needs to be compiled to create a binary program.
• Programs that require the presence of another programming language like Java, Python,
or Perl.
60
M. Kappelmann-Fenzl
Introduction
Many NGS data analysis programs are written in Java, Perl, C++ or Python. Simply
running these programs does not require programming knowledge, however, it is indeed
helpful. Within the framework of this textbook bioinformatic details are not covered.
Anyway, to be able to perform NGS data analysis familiarity with R is important. For
those of you who are not familiar with R, it is highly recommended to take advantage of
some freely available web resources (e.g. https://www.bioconductor.org/help/coursematerials/2012/SeattleMay2012/).
5.2
Computer Setup for NGS Data Analysis
Next-generation sequencing analysis is a computationally demanding process. Your average laptop is probably not up to the challenge. Any Linux/Unix-based operating system
will work well, with large servers in mind. Typically, analysis algorithms will be
distributed by researchers in one of the three ways:
• Standalone program to run in a Linux/Unix environment (most common).
• Webserver where you upload data to be analyzed (i.e. Galaxy).
• R software package/library for the R computing environment.
In theory, any computer with enough RAM, hard drive space, and CPU power can be
used for analysis. In general, you will need:
• 16 Gb of RAM Minimum (better to have 96+ Gb).
• 500 Gb of disk space (better to have 10+ Tb of space).
• Fast CPU (better to have at least 8 cores, more the better).
• External storage.
Depending on the type of analysis you want to perform, the required RAM, hard drive
space, and CPU power may vary and for some analyses a conventional laptop is fully
sufficient. If you own a computer or lab server with the aforementioned properties, you can
start over setting up your computer environment. The command line analysis tools that we
demonstrate in this book run on the Linux/Unix operating system. For practicing NGS data
analysis, we will use several software tools in various formats:
• Binary executable code that can be run directly on the target computer.
• Source code that needs to be compiled to create a binary program.
• Programs that require the presence of another programming language like Java, Python,
or Perl.
60
M. Kappelmann-Fenzl
