utilities with an evergrowing amount of data on their business operations and
infrastructure. Advanced metering technologies coupled with informatics create
an opportunity to form digital multiutility service providers [37]. Such metering
devices embrace two distinct technologies: meters that record water usage and
communication systems that can store and transmit real-time water use information
[38]. The ideal approach for their smart city application is installing smart water
meters at the property boundary in conjunction with intelligent end-use pattern
recognition algorithms either in-built into the meter software or within a processing
module at the utilities data centre. However, such an end goal requires the ability
to analyse collected data without human interaction and manual reclassification,
and this is non-trivial.
Increasing amounts of smart network data are now being collected by WSPs;
however the data is only of real business value if this valuable resource is ultimately
used to inform and support decision-making. The full range of uses for these
observations is only beginning to be realised and exploited. Recent work has
explored the use and analysis of such data. It has been argued that using
actual observed data, demand profiles can be calculated to provide more accurate
representations of high-granularity historical data, with potential applications in
real-time leakage detection, customer profiling and the provision of network
modelling demand patterns [39]. Time series clustering is an active area of research,
with the major issues being high dimensionality, temporal order and noise [40].
McKennaa et al. [41] investigated employing Gaussian mixture models (GMMs)
as the basis set for representing demand patterns using a dataset of hourly demand
readings spanning a 6-month study period, for 85 service connections within a
single DMA. While there was no customer information available for the dataset,
it was hypothesised after applying k-means that evidence of patterns found may
represent both residential and commercial customers. Garcia et al. [42] demonstrated
the potential use of k-means for clustering AMR data based on shape. Hadoop and
Spark were used in a big data context to provide an unsupervised classification
of the demand patterns from smart meters, with hourly interval feature vectors
of a weekly profile for 51,117 smart meters over a 1-year period (approximately
317 million observed readings). Nine distinctive clusters were identified. However,
no additional information (including customer type) was available other than
demands. In Mounce et al. [43], a case study of approximately 250 million readings
is presented, using a workflow for cleaning and preprocessing AMR data and
then clustering average daily demand patterns using the k-means ++ algorithm
with a correlation distance metric. Three natural clusters in the data (confirmed
by using silhouette plots) were found to correspond strongly to a residential and
commercial composition based on customer type which was used for post-analysis.
Further, a classification approach was also presented, comparing five classification
models with K-fold cross-validation, in order to classify into residential and
commercial customers. When using an ensemble of RUSBoosted decision trees
(for a fivefold), the overall accuracy was 91.3% (TPR 92% for residential and
84% for commercial) confirming dominant patterns of usage.
20
S. R. Mounce
Précédent

- 39/357

Suivant