379
2. Completeness may be straightforward to ascertain if there are empty fields and
gaps in the dataset. If we are lucky, the creators of the data left incomplete data
blank, rather than employing gap-filling techniques. Gap filling, especially if
done without disclosure, can seriously compromise data quality by misrepresenting
completeness and compromising validity.
3. Precision (and accuracy). The precision of valid data tends to be wellcharacterized so that we know how close the numbers are to reality. When mixing data of varying or unknown precision, it is important to accurately describe
the precision of the resulting derived data products or conclusions. Accuracy and
precision are closely related concepts via statistics.
4. Resolution. The resolution is the level of spatial, temporal, or categorical detail
at which a dataset is aggregated; for instance, the “energy” category involves
natural gas, gasoline, oil, electricity, etc., and “monthly” is a finer temporal resolution than “annual.” Resolution and scale are distinct but sometimes closely
correlated concepts; scale refers to the size of a process whereas resolution refers
to the detail of the data. If resolution is not finer than scale, there is a fundamental problem with the validity of the data (but this is common).
5. Provenance is the lineage and origin of the data. Auditability is the gold standard for provenance. However, this standard is rarely achieved for scientific
research data because this requires that the quality and provenance of data can be
verified by tracking it upstream through the data life cycle to its source, following a chain-of-custody; this auditability is achievable if each step in the data
chain (or supply chain) maintains records about its internal processes and also its
first-degree connections both upstream and downstream in the system. Do we
know where the data came from, what methods were used to process it, and what
the known problems are? Valid data tends to have solid provenance.
14.2.3 Data Scale and Resolution
Data Scale and Resolution are closely related concepts. Scale normally refers to the
dominant size, speed, generalizability, or frequency of a real-world process.
Resolution is the analogy to scale in the data domain. The two will sometimes be
used interchangeably in this text. There are some scales and resolutions of data and
process that are most relevant to FEW systems, and these are reviewed below. There
are four primary spatial resolutions or scales of data in FEW systems: Macro, Meso,
Establishment, and Process. These spatial scales allow us to describe agents, processes, and nodes in the FEW system’s network structure at some level of spatial
aggregation. All scales are useful for answering the same types of questions, but the
coarser data is less actionable, and the finer data is more expensive, less available,
and carries privacy and security implications. Transportation processes exist at all
scales; these processes are associated with the edges and flows between nodes in the
FEW system’s network graph structure.
14 Data
Précédent

- 386/686

Suivant