395
the national security sector… and then you must still demonstrate need-to-know. In
the USA, many National Laboratories have the necessary security credentials and
facilities to work with PCII and classified data, and are the primary venues for
research using these data. Also in the USA, the U.S. Census Bureau’s Research Data
Centers are secure facilities for research using PII data under proper data protection
protocols. Differential Privacy is a database technology that allows users to query
only those fields and joins that are permitted based on their credentials; this preserves the full potential utility of the data while tailoring privacy to suit the user’s
application. Differential privacy is a promising future tool to enable research, but it
is not yet broadly implemented for FEW systems data repositories.
Aggregation is a common technique for de-identifying PII or PCII or other private or sensitive data by reporting the space-time location, stock, flow, etc. for a
group instead of an individual. The coarser the resolution of the aggregation, the
greater the privacy and lower the utility and quality of the data. For utility customer
data, the Rule of Fifteen is a common best practice governing legally minimal
aggregation. The Rule of Fifteen is a common legal threshold in US State law
specifying that utility customer data and other PII data must be de-identified through
aggregation into groups of not less than 15 individuals, any one of which comprises
not more than 15% of the group’s total usage volume. This rule is a crude approximation of the census bureaus’ careful statistical practice of aggregating their PII
survey data at a resolution of census blocks, tracts, municipalities, or counties (etc.)
to preserve the privacy of individuals and companies’ data. The Rule of Fifteen is
easy to apply to residential utility customers, of which a city will have many thousands, but it difficult to apply to industrial and business customers, of which a city
may have only one or a few of a given type. An especially problematic type data
quality problem—incompleteness is created by census bureaus when especially
large companies are dropped from datasets for the sake of anonymization and privacy. Because the “fat tail” of the distribution of users—the extremely large producers and consumers—are responsible for a large fraction of the FEW system function,
they cannot be dropped without dramatically damaging the completeness and validity of the data. Thus aggregation is not an adequate solution to the privacy and
sensitivity problem, because it fails for some of the most important and valuable use
cases involving key industries or address-level analysis of energy and water usage.
The opposite technique, Disaggregation, attempts to reverse aggregation and
achieve finer resolution using assumptions which trade reduced validity for
increased precision/accuracy. Disaggregation is fundamentally a modeling
method that cannot create new information and is therefore only marginally useful—and should be used only with great caution to avoid misrepresentation. It is
far better to go get valid data at the resolution required, if possible than to estimate
those data using disaggregation. If this is not possible, one should ask themselves
whether it is ethical in this application to make inaccurate estimates of data that
are being intentionally held private for reasons of sensitivity. If the real data are
sensitive, then surely the estimated and inaccurate data are equally sensitive and
also potentially misleading.
14 Data
Précédent

- 402/686

Suivant