validation is limited to aggregated data (e.g. daily water use in a supply area). In
other cases, software tools are built to screen and flag the data which is identified as
suspicious.
Currently, it is not possible to validate all the variables. In general, the most
important data is validated, by a mix of manual and automatized routines. One of the
companies introduced the concept of a data diet, which implies a profound consideration of which data have to be measured, in which kind of time interval they have
to be stored and which of them have to be validated, before starting to generate data.
For some datasets, it is not really needed to develop a high level of validation. For
those datasets where validation is essential, we should look into the possibility of
correlation between different variables, measure all these variables and use correlation techniques (data science, statistics, models) for the validation.
The techniques used by the water companies to validate the data include:
• Manual validation (expert judgment)
• Visual comparison
• Control of measuring range, plausibility, data types
• Cross-correlation, statistical methods and models
• Combining own data with validated external sources
Examples of current practices on data validation are:
1. Filling missing data in the records of produced water using registered energy use
and relation between energy use and produced m
3 of water.
2. Determining missing year of installation of the pipes using the age of the
buildings of the area. Although these methods are not exact, they help to improve
the quality of the datasets. To the question regarding which platforms drinking
water companies use to store and process the data, each company has its own
(customized) systems. Some examples are shown in Table 1.
Table 1 Overview of tools per utility
Utility Systems
A
FEWS, Aspen and Midas
B
A MS SQL server database
C
PI (real-time information assets), SAP (context information of the assets) sample manager (information regarding water quality) and SCADA (events and all process
information)
D
PGIM (database van ABB 800x a process automatization) and own data warehouse
(Microsoft SQL)
E
PA (PIMS) and SAP
F
GIS (ESRI) own information system (accent) and SAP SharePoint
G
Data warehouse and Infor PGIM
H
Oracle Data warehouse, SQL, Excel, MS Power and BI ARCGIS
A Bird’s-Eye View of Data Validation in the Drinking Water Industry of the. . .
75
Précédent

- 93/357

Suivant