381
and climate, differences between suburbs and urban cores, and differences between
urban and rural economies.
In the USA, the USDA provides meso-scale survey data of uniformly high quality for crop production using standard national collection techniques with documented provenance, but the USGS provides meso-scale data for water use that is
aggregated from nonstandard local and state sources with quality and provenance
that varies dramatically from state to state. A good example of meso-scale data is
the county-level agricultural production dataset provided by the U.S. National
Agricultural Statistics Service (NASS 2015), which breaks out annual food production at the level of individual commodities like corn, or soybeans, beef, or the
Commodity Flow Survey that provides source and destination commodity transportation information between cities and counties in the USA (BTS 2012).
Aggregation methods at the meso-scale (like all aggregation) are often nonstandard and opaque, and disaggregated source data is by definition usually unavailable,
so meso-scale data cannot be quality controlled or audited if this is the case.
Establishing the provenance of meso-scale data, for this reason, is hard. Meso-scale
data accuracy and precision can be very good in one nation, region, or city, but very
poor in another. Additionally, the largest and most essential establishments (i.e., the
“fat tail” participants) are often dropped from meso-scale aggregations to preserve
business privacy, as in the case where one large company is responsible for the vast
majority of energy production in a local area. This is a significant hazard in mesoscale data products and can result in under-reporting of the most essential establishments in the FEW enterprise.
Some of the most important social, economic, and government interests in FEW
operate at the meso-scale, so meso-scale information is highly actionable especially
for local policy and infrastructure purposes. A large fraction, and probably the
majority, of the information content of bottom-up FEW data, is preserved at the
meso-scale. Meso-scale data is actionable for identifying infrastructure-scale
dependencies, stresses, and vulnerabilities, for managing city-scale FEW systems,
for linking cities with their rural FEW support systems, for water resource management, for mapping ecosystem-level hotspots in streams and terrestrial systems, for
making coarse-level sourcing and siting decisions, and for assessing vulnerabilities
at the scale of most human and natural disasters. Individual businesses are not
resolved, but clusters of FEW businesses are resolved at the meso-scale.
Meso-scale data is more computationally challenging to work with, but is still
relatively feasible given modern computing resources and software tools. Because
of the plurality of spatial units that encode meso-scale data, mismatched spatial
domains create significant problems for systems analysis.
The Establishment scale is often called the “address” scale or “customer” scale
of an individual building, residence, or facility. The establishment spatial scale may
correspond to any timescale of data. The Establishment scale is intrinsically private,
PCII, and/or PII categorized unless an exception is granted (e.g., for a public establishment that is not critical infrastructure). Establishments are often the resolution
at which data is reported due to the survey methods (mailed to specific addresses)
and sensors (e.g., water or gas meters) that are employed, but establishments tend to
14 Data
and climate, differences between suburbs and urban cores, and differences between
urban and rural economies.
In the USA, the USDA provides meso-scale survey data of uniformly high quality for crop production using standard national collection techniques with documented provenance, but the USGS provides meso-scale data for water use that is
aggregated from nonstandard local and state sources with quality and provenance
that varies dramatically from state to state. A good example of meso-scale data is
the county-level agricultural production dataset provided by the U.S. National
Agricultural Statistics Service (NASS 2015), which breaks out annual food production at the level of individual commodities like corn, or soybeans, beef, or the
Commodity Flow Survey that provides source and destination commodity transportation information between cities and counties in the USA (BTS 2012).
Aggregation methods at the meso-scale (like all aggregation) are often nonstandard and opaque, and disaggregated source data is by definition usually unavailable,
so meso-scale data cannot be quality controlled or audited if this is the case.
Establishing the provenance of meso-scale data, for this reason, is hard. Meso-scale
data accuracy and precision can be very good in one nation, region, or city, but very
poor in another. Additionally, the largest and most essential establishments (i.e., the
“fat tail” participants) are often dropped from meso-scale aggregations to preserve
business privacy, as in the case where one large company is responsible for the vast
majority of energy production in a local area. This is a significant hazard in mesoscale data products and can result in under-reporting of the most essential establishments in the FEW enterprise.
Some of the most important social, economic, and government interests in FEW
operate at the meso-scale, so meso-scale information is highly actionable especially
for local policy and infrastructure purposes. A large fraction, and probably the
majority, of the information content of bottom-up FEW data, is preserved at the
meso-scale. Meso-scale data is actionable for identifying infrastructure-scale
dependencies, stresses, and vulnerabilities, for managing city-scale FEW systems,
for linking cities with their rural FEW support systems, for water resource management, for mapping ecosystem-level hotspots in streams and terrestrial systems, for
making coarse-level sourcing and siting decisions, and for assessing vulnerabilities
at the scale of most human and natural disasters. Individual businesses are not
resolved, but clusters of FEW businesses are resolved at the meso-scale.
Meso-scale data is more computationally challenging to work with, but is still
relatively feasible given modern computing resources and software tools. Because
of the plurality of spatial units that encode meso-scale data, mismatched spatial
domains create significant problems for systems analysis.
The Establishment scale is often called the “address” scale or “customer” scale
of an individual building, residence, or facility. The establishment spatial scale may
correspond to any timescale of data. The Establishment scale is intrinsically private,
PCII, and/or PII categorized unless an exception is granted (e.g., for a public establishment that is not critical infrastructure). Establishments are often the resolution
at which data is reported due to the survey methods (mailed to specific addresses)
and sensors (e.g., water or gas meters) that are employed, but establishments tend to
14 Data
