the heart of the ability to derive actionable value from an array of structured and
increasingly unstructured, text based or sensory data and execute or automate the
next best action based on predictive and prescriptive data science.
The current nature of water utility network data is that it remains sparse
(e.g. not all locations are sampled) and typically is not linked across functions
(e.g. water quality data is not linked to hydraulic model data). Maximising
the quality of data (its usefulness) requires consideration of a chain of processes
and manipulation, e.g. data source, collection, storage and the anticipated data
end use. Machine learning or data-driven analyses, which map inputs to outputs
without attempting to accurately model underlying processes, can potentially yield
useful understanding, such as determination of dominant variables and empirical
relationships, and therefore have been used for many different environmental
and water quality applications. The required blend of foresight and experience
means a move towards ‘big data’ solutions, and so-called business intelligence
(turning an organisation’s data into patterns that help make intelligent business
decisions) in water utilities must be somewhat iterative and will require significant
development time. In the future, engineers will access data, tools and analysis in real
time and collaborate on diagnostic decisions relating to the condition of a remote
monitored asset. This could enable network engineering decisions to be increasingly
evidence-based with associated provenance to allow the reasoning behind decisions
to be evaluated for compliance with statutory regulation and to create a knowledge
repository to form a basis for future decisions. In an industry where customer
perception of service (and statutory obligations) is dependent on very complex,
distributed, non-linear dynamical water networks with a consequent high degree
of uncertainty, such knowledge-based engineering represents a key advantage
in commercial terms (i.e. company efficiency) and in the ability to serve the wider
society.
A vision for this integrated future is illustrated in Fig. 2. At the rate at which
data and our ability to analyse it are growing in society in general, it is reasonable
to expect that most UK companies will be using the impact of big data analytics
in the next 5 years [6]. It should however be noted that this is unlikely to be the
case in many other countries, with a more gradual trickle down of technology
transfer occurring over time.
Ultimately developments will result in data exploration tools for non-ICT
specialists reducing the cost of and ability for high-fidelity visualisation of data
for enabling human interpretation. The exploration and analysis of data using
visualisation techniques is a powerful approach for conveying potential hypotheses
and exploring correlations, due to the fact that vision plays an important role in
human cognition. One example of how data-driven techniques can be used for
data visualisation is the use of self-organising maps (SOM), a form of unsupervised
artificial neural networks (ANNs), as has been used for analysing water data
in various UK industry research projects [7]. The application of SOMs has been
demonstrated in water distribution system data mining for microbiological and
physico-chemical data at laboratory scale [8] and in the field [9]; in clustering of
water quality, hydraulic modelling and asset data for a single water supply zone [10];
6
S. R. Mounce
increasingly unstructured, text based or sensory data and execute or automate the
next best action based on predictive and prescriptive data science.
The current nature of water utility network data is that it remains sparse
(e.g. not all locations are sampled) and typically is not linked across functions
(e.g. water quality data is not linked to hydraulic model data). Maximising
the quality of data (its usefulness) requires consideration of a chain of processes
and manipulation, e.g. data source, collection, storage and the anticipated data
end use. Machine learning or data-driven analyses, which map inputs to outputs
without attempting to accurately model underlying processes, can potentially yield
useful understanding, such as determination of dominant variables and empirical
relationships, and therefore have been used for many different environmental
and water quality applications. The required blend of foresight and experience
means a move towards ‘big data’ solutions, and so-called business intelligence
(turning an organisation’s data into patterns that help make intelligent business
decisions) in water utilities must be somewhat iterative and will require significant
development time. In the future, engineers will access data, tools and analysis in real
time and collaborate on diagnostic decisions relating to the condition of a remote
monitored asset. This could enable network engineering decisions to be increasingly
evidence-based with associated provenance to allow the reasoning behind decisions
to be evaluated for compliance with statutory regulation and to create a knowledge
repository to form a basis for future decisions. In an industry where customer
perception of service (and statutory obligations) is dependent on very complex,
distributed, non-linear dynamical water networks with a consequent high degree
of uncertainty, such knowledge-based engineering represents a key advantage
in commercial terms (i.e. company efficiency) and in the ability to serve the wider
society.
A vision for this integrated future is illustrated in Fig. 2. At the rate at which
data and our ability to analyse it are growing in society in general, it is reasonable
to expect that most UK companies will be using the impact of big data analytics
in the next 5 years [6]. It should however be noted that this is unlikely to be the
case in many other countries, with a more gradual trickle down of technology
transfer occurring over time.
Ultimately developments will result in data exploration tools for non-ICT
specialists reducing the cost of and ability for high-fidelity visualisation of data
for enabling human interpretation. The exploration and analysis of data using
visualisation techniques is a powerful approach for conveying potential hypotheses
and exploring correlations, due to the fact that vision plays an important role in
human cognition. One example of how data-driven techniques can be used for
data visualisation is the use of self-organising maps (SOM), a form of unsupervised
artificial neural networks (ANNs), as has been used for analysing water data
in various UK industry research projects [7]. The application of SOMs has been
demonstrated in water distribution system data mining for microbiological and
physico-chemical data at laboratory scale [8] and in the field [9]; in clustering of
water quality, hydraulic modelling and asset data for a single water supply zone [10];
6
S. R. Mounce
