90
3 Jet Substructure at the LHC
a loss function. The free parameters of the algorithms are determined with training
sets, which is known as supervised learning. It has been found that in many cases,
adding more than the two or three variables with highest significance to the tagging
algorithm does not result in a significant performance gain. In these studies, also
the linear correlations of substructure variables have been examined, indicating the
mutual information carried by observables. No new observables linearly uncorrelated
to the output distribution could be identified, which would have indicated a potential
gain in performance when adding these to the new tagger.
New developments in machine learning (ML) allow for more complex architectures of artificial neural networks (ANNs), with several hidden layers and a large
number of neurons in each layer—so called deep neural networks (DNNs). These
networks are capable of learning complex, non-linear correlations when trained
on a sufficiently large sample of simulated data [536]. While previous approaches
like the ones mentioned above have combined shallow ANNs with high-level observables, DNNs allow to obtain the classification power directly from the raw data using
low-level inputs such as the four-vectors of all reconstructed jet constituents or just
energy deposits in the detector. A further advantage of DNNs is the possibility to
have multi-class classification with one output per class, instead of only discriminating between background and signal. There are several representations conceivable
which can be used to pass the full low-level information to DNNs. Jet images are pixelated images, where the pixel intensity represents the momentum of all particles that
deposited energy in a particular angular region [537–540]. Additional information,
such as the identified particle type, can be encoded by additional image layers, similar
to colour images [477]. Jet images can be analysed with convoluted neural networks
(CNNs), which have a smaller number of parameters to be determined in the training
than recurrent neural networks (RNNs). These RNNs can be used to analyse sequential information, such as the four-vectors of all jet constituents ordered in p T [541].
This approach has already been been used for jet flavour tagging (b or c tagging),
with charged particle tracks and secondary vertices as inputs [504, 542–544]. A
generalisation of this approach, where jets are represented as graphs, has been studied in the context of pileup mitigation [545]. Lastly, advances in ML tools on point
clouds have allowed the study of jets as unordered sets with the possibility to encode
additional information about particles beyond their kinematic properties [225, 546].
More information on these ML tools in the context of jet substructure can be found
in a recent review [24].
Recently, advanced machine learning methods have been applied to jet substructure taggers by ATLAS [525] and CMS [526]. In a first step, ATLAS has used 10–13
high-level observables to train a BDT and a DNN for W and t tagging. It has been
found that the BDT and DNN performances are nearly identical, indicating that
the advanced DNN can not surpass the BDT performance through an algorithmic
improvement if the inputs are the same.
11 A comparison of the ROC curves of these
two taggers (BDT top and DNN top) is shown in Fig. 3.13 (left). The ML taggers
11 A similar conclusion has been obtained in a comprehensive comparison of DNN-based top taggers,
where only small performance differences were observed due to differences in the algorithms [547].
Précédent

- 104/298

Suivant