377
research. The U.S. Department of Agriculture maintains a system of codings for
agricultural production associated with the National Agricultural Statistics Service
(NASS 2015). The U.S. Geological Survey maintains a system of codings for
water use in its Water Census of the USA (Maupin et al. 2014). The U.S. Energy
Information administration similarly classifies energy production and use (EIA
2017). There are many other examples of data codings that apply to a single dataset and layer of the FEW system—but they do not form complete FEW systems
ontologies. The FEWSION project has undertaken a relatively complete but not
particularly detailed ontology for the FEW system (FEWSION, n.d.; FEWSION
Codesheet 1.0, 2019).
Codings created for disparate purposes often do not line up one-to-one between
different layers of the FEW system network. These datasets are “heterogeneous,”
meaning that they are not a part of a consistent and coherent ontology. As a result, a
common research task in FEW systems is the development of FEW system ontologies that relate one dataset to another. Unfortunately, at this early stage of FEW
systems research there has not been sufficient progress on the creating of standardized ontologies including the relational data models, controlled vocabularies, and
crosswalks, and as a result, most of this work currently involves custom and
application- specific bilateral crosswalks. There is currently a pressing need for the
development of standard FEW systems ontologies.
In the absence of standard FEW systems ontologies, there are a number of problems that need to be solved in order to pair up two datasets. When more than two
datasets need to be linked, it is necessary to either construct a complete relational
data model or to choose a single lowest common denominator (LCD) dataset to
which all other datasets will be related in the ontology. For instance, if several datasets and models have different resolution, one can standardize them at the coarsest
resolution of space, time, and category. This is roughly the approach taken by the
FEWSION project in its selection of the county, month, and FEWSION/SCTG+
category codes (FEWSION Codesheet 1.0, 2019; BTS, n.d.); this resolution is more
or less compatible with all of the available data comprising the systems network
description. Building a robust and widely encompassing FEW systems ontology is
a significant undertaking that has not yet been completed, although various input–
output and commodity flow modeling teams have made strides especially at the
coarser resolutions. The LCD approach is a simpler but also less flexible solution.
The steps involved in the more common LCD approach commonly include, but are
not limited to:
1. Identifying a formal coding and controlled vocabulary for each dataset.
2. Establishing a crosswalk between each pair of controlled vocabularies.
3. Developing aggregation or disaggregation factors to relate datasets with mismatching resolution.
Data Format is the most common colloquial shorthand for a data structure in
FEW science and management work. Most communities have de-facto best practices for data structure encoded in their preferred data formats. For instance, many
14 Data
research. The U.S. Department of Agriculture maintains a system of codings for
agricultural production associated with the National Agricultural Statistics Service
(NASS 2015). The U.S. Geological Survey maintains a system of codings for
water use in its Water Census of the USA (Maupin et al. 2014). The U.S. Energy
Information administration similarly classifies energy production and use (EIA
2017). There are many other examples of data codings that apply to a single dataset and layer of the FEW system—but they do not form complete FEW systems
ontologies. The FEWSION project has undertaken a relatively complete but not
particularly detailed ontology for the FEW system (FEWSION, n.d.; FEWSION
Codesheet 1.0, 2019).
Codings created for disparate purposes often do not line up one-to-one between
different layers of the FEW system network. These datasets are “heterogeneous,”
meaning that they are not a part of a consistent and coherent ontology. As a result, a
common research task in FEW systems is the development of FEW system ontologies that relate one dataset to another. Unfortunately, at this early stage of FEW
systems research there has not been sufficient progress on the creating of standardized ontologies including the relational data models, controlled vocabularies, and
crosswalks, and as a result, most of this work currently involves custom and
application- specific bilateral crosswalks. There is currently a pressing need for the
development of standard FEW systems ontologies.
In the absence of standard FEW systems ontologies, there are a number of problems that need to be solved in order to pair up two datasets. When more than two
datasets need to be linked, it is necessary to either construct a complete relational
data model or to choose a single lowest common denominator (LCD) dataset to
which all other datasets will be related in the ontology. For instance, if several datasets and models have different resolution, one can standardize them at the coarsest
resolution of space, time, and category. This is roughly the approach taken by the
FEWSION project in its selection of the county, month, and FEWSION/SCTG+
category codes (FEWSION Codesheet 1.0, 2019; BTS, n.d.); this resolution is more
or less compatible with all of the available data comprising the systems network
description. Building a robust and widely encompassing FEW systems ontology is
a significant undertaking that has not yet been completed, although various input–
output and commodity flow modeling teams have made strides especially at the
coarser resolutions. The LCD approach is a simpler but also less flexible solution.
The steps involved in the more common LCD approach commonly include, but are
not limited to:
1. Identifying a formal coding and controlled vocabulary for each dataset.
2. Establishing a crosswalk between each pair of controlled vocabularies.
3. Developing aggregation or disaggregation factors to relate datasets with mismatching resolution.
Data Format is the most common colloquial shorthand for a data structure in
FEW science and management work. Most communities have de-facto best practices for data structure encoded in their preferred data formats. For instance, many
14 Data
