398
• Rules for privacy, protection, ethics, and also laws governing individually
identifiable and critical infrastructure data are overriding concerns for FEW
systems data.
• Some common data aggregation and access principles explain the type of FEW
systems data that is commonly accessible (and nonaccessible) in the USA.
• There are common data scales at which FEW systems are described, and these
data scales do not match all problem scales; high-quality FEW systems data is
accessible for some geographies, scales, and application domains, and not for
others.
• FEW systems data cover a myriad of highly specialized public and private applications, and these are voluminous, complex, and diverse with respect to data
structure and standard, as well as the repositories that handle each application.
Discussion Points and Exercises
1. Many types of data are only useful when combined with other data that may not
yet exist. The value of these data grows over time. For example, by “joining”
independently collected energy use, water use, demographic, building-type,
and process level data, it might be possible for us to precisely understand how
to help make a given business, household, or process more efficient and sustainable. What if this exact type of combined use of the data was not anticipated at
the time of its collection several decades earlier, and the person or business
about whom the data was originally collected no longer exists to seek consent?
Is it ethical to make this use of the data, or not?
2. Have you ever had your privacy violated? How did it feel?
3. Is there a specific type of food, energy, or water data about yourself or your
business that you would not be comfortable sharing with the world? Why or
why not? And, at what level of aggregation of that data with others’ data would
you become comfortable with the release? Why?
4. Identify an example of data “security through obscurity” in the FEW data
space, and distinguish that example clearly from a similar case where data protection was intentionally applied to restrict access to data.
5. Is there a specific type of FEW data that you need for your work and have been
unable to access? Why was that—data protection practices, law, privacy, or
poor application of FAIR principles?
6. Explain the four major FAIR data management principles.
7. Choose a data repository with which you are familiar and evaluate it against
both the FAIR data management principles and data ethics principles.
8. Identify an appropriate repository for a dataset that you would like to publish,
following the COPDESS and/or Stanford Library guidelines. Attempt to follow
that repository’s processes to publish your data.
9. Identify the steps in the Data Life Cycle, in your own words.
10. Have you ever been unable to find a dataset? Why and what was the specific
problem?
11. Describe how data aggregation works and precisely identify the difference
between adequate and inadequate levels of data aggregation for a private or
sensitive FEW dataset with which you are familiar.
B. L. Ruddell
• Rules for privacy, protection, ethics, and also laws governing individually
identifiable and critical infrastructure data are overriding concerns for FEW
systems data.
• Some common data aggregation and access principles explain the type of FEW
systems data that is commonly accessible (and nonaccessible) in the USA.
• There are common data scales at which FEW systems are described, and these
data scales do not match all problem scales; high-quality FEW systems data is
accessible for some geographies, scales, and application domains, and not for
others.
• FEW systems data cover a myriad of highly specialized public and private applications, and these are voluminous, complex, and diverse with respect to data
structure and standard, as well as the repositories that handle each application.
Discussion Points and Exercises
1. Many types of data are only useful when combined with other data that may not
yet exist. The value of these data grows over time. For example, by “joining”
independently collected energy use, water use, demographic, building-type,
and process level data, it might be possible for us to precisely understand how
to help make a given business, household, or process more efficient and sustainable. What if this exact type of combined use of the data was not anticipated at
the time of its collection several decades earlier, and the person or business
about whom the data was originally collected no longer exists to seek consent?
Is it ethical to make this use of the data, or not?
2. Have you ever had your privacy violated? How did it feel?
3. Is there a specific type of food, energy, or water data about yourself or your
business that you would not be comfortable sharing with the world? Why or
why not? And, at what level of aggregation of that data with others’ data would
you become comfortable with the release? Why?
4. Identify an example of data “security through obscurity” in the FEW data
space, and distinguish that example clearly from a similar case where data protection was intentionally applied to restrict access to data.
5. Is there a specific type of FEW data that you need for your work and have been
unable to access? Why was that—data protection practices, law, privacy, or
poor application of FAIR principles?
6. Explain the four major FAIR data management principles.
7. Choose a data repository with which you are familiar and evaluate it against
both the FAIR data management principles and data ethics principles.
8. Identify an appropriate repository for a dataset that you would like to publish,
following the COPDESS and/or Stanford Library guidelines. Attempt to follow
that repository’s processes to publish your data.
9. Identify the steps in the Data Life Cycle, in your own words.
10. Have you ever been unable to find a dataset? Why and what was the specific
problem?
11. Describe how data aggregation works and precisely identify the difference
between adequate and inadequate levels of data aggregation for a private or
sensitive FEW dataset with which you are familiar.
B. L. Ruddell
