IN SITU OBSERVATIONS: SYSTEMS AND MANAGEMENT
225
with other datasets relevant for his application. They must help to find the
data among the network (data catalogues). That is the purpose of defining
correctly distribution data format as well as the metadata (data on the data)
that need to be preserved for future processing.
Data format have always been a nightmare both for users and data
managers and they are both dreaming of the "Esperanto" of data format.
Computer technology has improved a lot in the past decade and we are
slowly moving from ASCII format (easy to use by human eyes but not for
softwares), to binary format (easy for software but not shareable among
platforms (Windows, Unix, etc), and self-descriptive, multiplatform formats
(Netcdf, Hdf, etc) that allow more flexibility in sharing data among a
network and are read by all softwares that are commonly used by scientists.
Depending on who is using the oceanographic data, the information
stored in a dataset can be more or less precise. When a scientist is using data
that he has acquired himself on a cruise, he has a lot of additional
information (often in his head) and he is mainly interested by the
measurements themselves. When he starts to share with other persons from
his laboratory he has to tell them how he took the measurements, from which
platform, what the sea-state was that day, what are the corrections he applied
on the raw data, etc in order for his colleagues to use the data properly and
understand differences with other datasets. When these data are made
available to a larger community the number of necessary additional
information, to be stored with the data themselves, increase, especially when
climatological or long-run re-analysis are some of the targeted applications.
This is why nowadays a lot of metadata are attached to any data shared
among a community.
One important point for metadata is to identify a common vocabulary to
record most of these information. This is pretty easy to achieve for a specific
community such as ARGO, but it starts to be a bit more difficult when we
want to address multidisciplinary datasets such as mooring data. To help
community in this area some metadata standards are emerging for the marine
community with Marine XML under ICES/IOC umbrella and ISO19115
norm.
Another important point in data format is to keep, together with the data,
the history of the processing and corrections that have been applied to it.
This is the purpose to the history-records that track what happened and allow
going back to data centres to ask for a previous version if a user wants to
perform his own processing from an earlier stage.
225
with other datasets relevant for his application. They must help to find the
data among the network (data catalogues). That is the purpose of defining
correctly distribution data format as well as the metadata (data on the data)
that need to be preserved for future processing.
Data format have always been a nightmare both for users and data
managers and they are both dreaming of the "Esperanto" of data format.
Computer technology has improved a lot in the past decade and we are
slowly moving from ASCII format (easy to use by human eyes but not for
softwares), to binary format (easy for software but not shareable among
platforms (Windows, Unix, etc), and self-descriptive, multiplatform formats
(Netcdf, Hdf, etc) that allow more flexibility in sharing data among a
network and are read by all softwares that are commonly used by scientists.
Depending on who is using the oceanographic data, the information
stored in a dataset can be more or less precise. When a scientist is using data
that he has acquired himself on a cruise, he has a lot of additional
information (often in his head) and he is mainly interested by the
measurements themselves. When he starts to share with other persons from
his laboratory he has to tell them how he took the measurements, from which
platform, what the sea-state was that day, what are the corrections he applied
on the raw data, etc in order for his colleagues to use the data properly and
understand differences with other datasets. When these data are made
available to a larger community the number of necessary additional
information, to be stored with the data themselves, increase, especially when
climatological or long-run re-analysis are some of the targeted applications.
This is why nowadays a lot of metadata are attached to any data shared
among a community.
One important point for metadata is to identify a common vocabulary to
record most of these information. This is pretty easy to achieve for a specific
community such as ARGO, but it starts to be a bit more difficult when we
want to address multidisciplinary datasets such as mooring data. To help
community in this area some metadata standards are emerging for the marine
community with Marine XML under ICES/IOC umbrella and ISO19115
norm.
Another important point in data format is to keep, together with the data,
the history of the processing and corrections that have been applied to it.
This is the purpose to the history-records that track what happened and allow
going back to data centres to ask for a previous version if a user wants to
perform his own processing from an earlier stage.
