5.4 Toward an Open Molecular Science
183
partners are regularly collected and analyzed in order to be adequately prioritized to
the end of maximizing the interoperability with the existing e-infrastructures.
5.4.3 Compute Resources and Data Management for
Molecular Open Science
Computational activities of the molecular science community rely on the usage of the
EGI Federated Cloud and PRACE resources. These compute resources are complemented by those of the distributed European Infrastructures and vocational national
and regional facilities (ranging from multicore processors to cloud clusters) among
which the most important are as follows:
CMAST a Virtual Laboratory to support research in chemistry integrated into the
computing infrastructure CRESCO (a production grid of computational resources
belonging to ENEA DTE-ICT).
RECAS a consortial of four national Italian data centers (Napoli, Bari, Catania,
Cosenza) that is part of the European Grid Infrastructure (EGI) and INFN.
MOSGRID a compute infrastructure that provides Grid services for molecular
simulations leveraging on an extensive use of the German D-Grid-Infrastructure
for high-performance computing to handle metadata and provide data mining and
knowledge generation.
Leveraging on these resources and on existing tools (see also the next subsection)
an Open Science Data Cloud (OSDC) can be established for the molecular science
community. The OSDC allows scientists to manage, analyze, share, and archive their
datasets. Datasets can be downloaded from the OSDC by anyone. The OSDC is not
only designed to provide a long-term persistent home for scientific data, but also
to provide a platform for data-intensive science so that new types of data-intensive
algorithms can be developed, tested, and used over large amounts of heterogeneous
scientific data. The not-for-profit OSDC is articulated as follows:
(1) Use a community of users and data curators to identify data to add to OSDC.
(2) Use permanent IDs to identify this data and associate metadata with these IDs.
(3) Support permissions so that colleagues can access this data prior to its public
release and to support analysis of access-controlled data.
(4) Support both file-based descriptors and APIs to access the data.
(5) Make available computing images via infrastructure as a service that contains
the software tools and applications commonly used by a community.
(6) Provide mechanisms to both import and export data and the associated computing environment so that researchers can easily move their computing infrastructures
between science clouds.
(7) Identify a sustainable level of investment in computing infrastructure and
operations and invest this amount each year.
(8) Provide general support for a limited number of applications.
183
partners are regularly collected and analyzed in order to be adequately prioritized to
the end of maximizing the interoperability with the existing e-infrastructures.
5.4.3 Compute Resources and Data Management for
Molecular Open Science
Computational activities of the molecular science community rely on the usage of the
EGI Federated Cloud and PRACE resources. These compute resources are complemented by those of the distributed European Infrastructures and vocational national
and regional facilities (ranging from multicore processors to cloud clusters) among
which the most important are as follows:
CMAST a Virtual Laboratory to support research in chemistry integrated into the
computing infrastructure CRESCO (a production grid of computational resources
belonging to ENEA DTE-ICT).
RECAS a consortial of four national Italian data centers (Napoli, Bari, Catania,
Cosenza) that is part of the European Grid Infrastructure (EGI) and INFN.
MOSGRID a compute infrastructure that provides Grid services for molecular
simulations leveraging on an extensive use of the German D-Grid-Infrastructure
for high-performance computing to handle metadata and provide data mining and
knowledge generation.
Leveraging on these resources and on existing tools (see also the next subsection)
an Open Science Data Cloud (OSDC) can be established for the molecular science
community. The OSDC allows scientists to manage, analyze, share, and archive their
datasets. Datasets can be downloaded from the OSDC by anyone. The OSDC is not
only designed to provide a long-term persistent home for scientific data, but also
to provide a platform for data-intensive science so that new types of data-intensive
algorithms can be developed, tested, and used over large amounts of heterogeneous
scientific data. The not-for-profit OSDC is articulated as follows:
(1) Use a community of users and data curators to identify data to add to OSDC.
(2) Use permanent IDs to identify this data and associate metadata with these IDs.
(3) Support permissions so that colleagues can access this data prior to its public
release and to support analysis of access-controlled data.
(4) Support both file-based descriptors and APIs to access the data.
(5) Make available computing images via infrastructure as a service that contains
the software tools and applications commonly used by a community.
(6) Provide mechanisms to both import and export data and the associated computing environment so that researchers can easily move their computing infrastructures
between science clouds.
(7) Identify a sustainable level of investment in computing infrastructure and
operations and invest this amount each year.
(8) Provide general support for a limited number of applications.
