387
14.3 FAIR Data Management and the Data Life Cycle
14.3.1 The Data Life Cycle
The Data Life Cycle involves experimental and observational design and approval,
data collection, quality control, metadata description, data curation with a repository, discovery (if necessary) of the data resource, integration of the data with other
data, and analysis of the data to answer questions. Sound data management practices are essential to enable the full completion (and repetition) of the data life cycle.
Data grows more valuable as it proceeds along its life cycle, and even more so with
reuse and repeated integration. You should make sure to implement sound data
management practices with any data that you collect so that it can benefit yourself
and others. (Re)using a scientific workflow or utilizing a mature data repository’s
workflow systems for data processing can help a great deal.
Unfortunately, we are still very early in the information age, and we have little
successful experience with the implementation of the data life cycle to date.
Common struggles include critically flawed metadata, use of nonstandard data formats, the lack of use of a formal and well documented ontology or controlled vocabulary, archival on websites or outside standard repositories where findability and
accessibility are poor, poor and irrecoverable initial quality control on data and
metadata, and lack of issuance of a unique identifier (DOI) for data (Ruddell, 2006).
Most critically, the “half-life” problem rapidly erodes our data resources as poor
data management processes result in datasets and/or critical metadata or the tools to
access those data vanishing steadily over time. The half-life of research data has
been estimated at under 10 years; this is an indefensible tragedy, given the expense
and value of data in a digital age (Ruddell et al., 2014).
14.3.2 FAIR Data Management
The FAIR Principles for scientific data management and stewardship were published in 2016, and provide guidelines for improving findability, accessibility,
interoperability, and reuse of data. Findability emphasizes the use of a globally
unique identifier for the dataset, use of that identifier within a rich metadata file that
is easy for both humans and computers to use, and the metadata is registered with
an appropriate searchable index service.
Accessibility emphasizes that data and metadata are retrievable using a standard
and open and free communication protocol, and that metadata accessibility persists
even when data are not accessible. Interoperability emphasizes that metadata
and data:
• Use a formal and widely used language or format.
• Use controlled vocabularies that follow FAIR principles.
• Appropriately reference other FAIR metadata and data.
14 Data
14.3 FAIR Data Management and the Data Life Cycle
14.3.1 The Data Life Cycle
The Data Life Cycle involves experimental and observational design and approval,
data collection, quality control, metadata description, data curation with a repository, discovery (if necessary) of the data resource, integration of the data with other
data, and analysis of the data to answer questions. Sound data management practices are essential to enable the full completion (and repetition) of the data life cycle.
Data grows more valuable as it proceeds along its life cycle, and even more so with
reuse and repeated integration. You should make sure to implement sound data
management practices with any data that you collect so that it can benefit yourself
and others. (Re)using a scientific workflow or utilizing a mature data repository’s
workflow systems for data processing can help a great deal.
Unfortunately, we are still very early in the information age, and we have little
successful experience with the implementation of the data life cycle to date.
Common struggles include critically flawed metadata, use of nonstandard data formats, the lack of use of a formal and well documented ontology or controlled vocabulary, archival on websites or outside standard repositories where findability and
accessibility are poor, poor and irrecoverable initial quality control on data and
metadata, and lack of issuance of a unique identifier (DOI) for data (Ruddell, 2006).
Most critically, the “half-life” problem rapidly erodes our data resources as poor
data management processes result in datasets and/or critical metadata or the tools to
access those data vanishing steadily over time. The half-life of research data has
been estimated at under 10 years; this is an indefensible tragedy, given the expense
and value of data in a digital age (Ruddell et al., 2014).
14.3.2 FAIR Data Management
The FAIR Principles for scientific data management and stewardship were published in 2016, and provide guidelines for improving findability, accessibility,
interoperability, and reuse of data. Findability emphasizes the use of a globally
unique identifier for the dataset, use of that identifier within a rich metadata file that
is easy for both humans and computers to use, and the metadata is registered with
an appropriate searchable index service.
Accessibility emphasizes that data and metadata are retrievable using a standard
and open and free communication protocol, and that metadata accessibility persists
even when data are not accessible. Interoperability emphasizes that metadata
and data:
• Use a formal and widely used language or format.
• Use controlled vocabularies that follow FAIR principles.
• Appropriately reference other FAIR metadata and data.
14 Data
