102
to label samples with these classes. Two different data types were used, the geochemical assays
and the manually identified material groups present in each sample.
2 BACKGROUND
2.1 K-means
K-means clustering involves separating n observations into k clusters to minimise the withincluster sum of squares (WCSS) (Steinhaus 1957). At the same time the between-cluster sum
of squares (BCSS) should be maximised. In this study, Squared Euclidean distance was used
to find the distance between data points and every centroid, where every centroid is the random number of n samples. The whole process of k-means is calculating the distances between
different data repeatedly to find out the best classification that minimises the WCSS and
maximises the BCSS in unsupervised method.
2.2 SVM
Support vector machine (SVM) (Melgani & Bruzzone 2004; Liu et al. 2016) can be used for
supervised classification and regression. SVM uses the construction of hyperplanes in highdimensional or infinite-dimensional space and a functional margin that indicates the distance
from hyperplane to the nearest training-data point of any class. When training the SVM the
functional margin should be maximised and the generalization error of the classifier minimised. SVM uses a kernel function to define dot products in order to simplify the computation in the original space. The formula: ∑ i α i k(x i ,x) = c relates the points x in the feature space
where c is a constant. This means the relative closeness of test points and data points can be
calculated by the sum of this kernel function.
2.3 Deep learning using autoencoder
An autoencoder is a type of neural network that can be used as an method for training a deep
learning neural network. Autoencoders reduce the dimension of the inputs to lower-dimensional code and then produce new features that have same number of input features in an
unsupervised manner (Bengio 2009). The learning is done by using 3 components: encoder,
code and decoder. The encoder compresses the input and produces the code, the decoder
then reconstructs the input only using this code. For learning new features autoencoders do
not use explicit training labels, thus the process can be considered an unsupervised learning
technique. However, as autoencoders produce their own labels from the input samples this is
generally considered as self-supervised or semi-supervised method.
In this study, two autoencoders are used for the deep learning. The features extracted
from the hidden layer of the first autoencoder are used as inputs to the second autoencoder.
Extracted features from the second autoencoder are used to train a classification layer with
the corresponding class labels. Therefore the deep net was created by stacking two autoencoders and a classification layer. Then the deep net was used to predict the labels for the test
samples.
2.4 Data
The data used in this paper is available through the Western Australian Department of Mines
and Petroleum (WAMEX A95838 2012) and has been previously studied in Nathan et al.
(2017). Drill hole data was obtained from 11 exploration holes, with a total depth of 662 m.
The data includes geochemical assays (Fe, Al 2 O 3 , P, SiO 2 , Mn, MgO, K 2 O, TiO 2 , S, As, Co,
Cu, Pb, Ni, Zn, Ba, Cr, Sn, V, Cl and loss on ignition (LOI)) and mineralogy (grouped into
BIF, shale, goethite and hematite) sampled at 2m intervals. Each of these intervals had a
manual classification for the ground truth.
Précédent

- 123/780

Suivant