378
climate models now use a specific kind of NetCDF implementation of the HDF5
format to capture the voluminous and multi-type gridded data that is produced by
these models. This practice is not mandated by any regulatory body, but it has been
adopted by the community and is, therefore, a de-facto standard. There are not de-facto
standards yet for the young FEW systems community or related fields of footprinting, sustainability analysis, life cycle analysis, agent-based modeling, and network
analysis. The reader is directed to the various repositories listed later in this chapter,
and in the end matter, for examples of de-facto data structure standards.
A common practice is the use of “flat” CSV text files and spreadsheets for smaller
volumes of data, with an accompanying “header” that lists the names and units of
each column of data, and a separate documentation file that defines terms. Often
FEW data text files are pseudo-spatial in that they are cross-coded and joined to
spatial locations such as rivers, counties, states, etc. that can be mapped with a
Geographic Information System. Remote sensing datasets for FEW topics have
their own well-developed format standards that are not unique to the FEW systems
domain of applications.
Learning to make use of these formats and de-facto standards is one of the core
outcomes and skillsets resulting from a research apprenticeship or early career
training experience within a specific domain of application. There are no general
rules—only best practices within your application domain. Learn to carefully identify which community of practice you will be working within and interoperating
with, and to ask experienced experts in that community the right questions about
what is the best, easiest, and most well-tested set of practices for data structures and
formats. It is inadvisable to become an early adopter who tries to introduce new
technologies, standards, and data structures to a community before you become an
expert on the older ways. A little bit of due diligence at the early stage of your work
will save you a great deal of time and yield much more rapid productivity. Data
format and structure is like a language; none is right, all have their place and
intended use, and it is most important that everyone is “on the same page.”
14.2.2 Data Quality
The major issues in FEW systems data quality are validity, completeness, precision
(and accuracy), resolution, and provenance. We need sufficient data quality to do
effective research and make effective operational decisions. Quality is entirely relative to the intended use.
1. Validity is the idea that data demonstrably corresponds to reality. Validity must
be rigorously ensured by using formal methods, carefully applied, in data collection. Because most FEW systems datasets are derived from statistical sampling
and surveys, it is important that those survey samples are unbiased and have a
large enough sample size. The design of the survey itself must be rigorous and
studied, with verifiable performance.
B. L. Ruddell
climate models now use a specific kind of NetCDF implementation of the HDF5
format to capture the voluminous and multi-type gridded data that is produced by
these models. This practice is not mandated by any regulatory body, but it has been
adopted by the community and is, therefore, a de-facto standard. There are not de-facto
standards yet for the young FEW systems community or related fields of footprinting, sustainability analysis, life cycle analysis, agent-based modeling, and network
analysis. The reader is directed to the various repositories listed later in this chapter,
and in the end matter, for examples of de-facto data structure standards.
A common practice is the use of “flat” CSV text files and spreadsheets for smaller
volumes of data, with an accompanying “header” that lists the names and units of
each column of data, and a separate documentation file that defines terms. Often
FEW data text files are pseudo-spatial in that they are cross-coded and joined to
spatial locations such as rivers, counties, states, etc. that can be mapped with a
Geographic Information System. Remote sensing datasets for FEW topics have
their own well-developed format standards that are not unique to the FEW systems
domain of applications.
Learning to make use of these formats and de-facto standards is one of the core
outcomes and skillsets resulting from a research apprenticeship or early career
training experience within a specific domain of application. There are no general
rules—only best practices within your application domain. Learn to carefully identify which community of practice you will be working within and interoperating
with, and to ask experienced experts in that community the right questions about
what is the best, easiest, and most well-tested set of practices for data structures and
formats. It is inadvisable to become an early adopter who tries to introduce new
technologies, standards, and data structures to a community before you become an
expert on the older ways. A little bit of due diligence at the early stage of your work
will save you a great deal of time and yield much more rapid productivity. Data
format and structure is like a language; none is right, all have their place and
intended use, and it is most important that everyone is “on the same page.”
14.2.2 Data Quality
The major issues in FEW systems data quality are validity, completeness, precision
(and accuracy), resolution, and provenance. We need sufficient data quality to do
effective research and make effective operational decisions. Quality is entirely relative to the intended use.
1. Validity is the idea that data demonstrably corresponds to reality. Validity must
be rigorously ensured by using formal methods, carefully applied, in data collection. Because most FEW systems datasets are derived from statistical sampling
and surveys, it is important that those survey samples are unbiased and have a
large enough sample size. The design of the survey itself must be rigorous and
studied, with verifiable performance.
B. L. Ruddell
