88
3 Jet Substructure at the LHC
approximately 20%. There exist other approaches to cover the transition from low
to high Lorentz boosts, using a single algorithm. In the HTTv2 algorithm, the jet
size is reduced until an optimal size R opt is found, defined by the fractional jet mass
contained in the smaller jet. This results in better performance at high p T , while keeping a low misidentification rate at low p T . Similar performance is obtained with the
HOTVR algorithm, although with less algorithmic complexity. A comparison of the
efficiencies and misidentification rates for the CMSTT, HTT, HTTv2 and HOTVR
as a function of p T is shown in Fig. 3.11 (right). This comparison has been obtained
with the parameter settings described in [245]. The tagger working points have been
chosen to obtain an efficiency of 30% at p T = 700 GeV. The turn-on of the HOTVR
algorithm is not as sharp as for the HTT or HTTv2 taggers, but the misidentification
rate is smaller in this region. Overall, the HOTVR tagger achieves a misidentification
rate comparable to the one of the CMSTT over the full range in p T , while achieving
similar efficiencies as the HTT or HTTv2 taggers at low p T . The high efficiency
of the HOTVR tagger starting from p T > 300 GeV, together with a flat background
efficiency, will make this tagger an important tool in analyses targeting intermediate
to high top quark boosts.
An important step towards the commissioning of top taggers within an experiment
are measurements of the efficiency and misidentification rate in real collision data.
Generally, high-purity samples of top jets in data are obtained using a tight signal
selection (a lepton, well-separated from a high- p T large-R jet, and an additional btagged jet) to ensure that events contain a fully-merged top quark decay in a single
large-R jet. This can never be fully achieved, as no requirements on the substructure of the large-R jet can be imposed without biasing the efficiency measurement.
This results in an efficiency measurement that will be based on a sample containing
partially-merged and non-merged top quark decays. These can be subtracted from
the efficiency measurement by using simulated events, as done in a measurement
by ATLAS [520]. This has the drawback of relying on a specific simulation and
the ambiguous definition of a fully merged top quark decay. By not correcting for
non-merged top quark decays, efficiency values are obtained smaller than the ones
suggested by ROC curve studies, see for example [459]. Instead of subtracting the
t backgrounds, recent measurements perform a simultaneous extraction of the efficiencies for fully and partially merged categories [525, 526, 530, 531]. This is done
by separating events into a passing and a failing region, where pass and fail are determined by a selection on τ 32 . The obtained distributions in the (groomed) jet mass are
used for a template fit of the fully merged, semi-merged and un-merged categories
to data, as shown in Fig. 3.12. The jet mass is used since the three contributions
have different shapes in this distribution, allowing for a precise determination of the
top tagging efficiencies. An additional advantage is that systematic uncertainties can
be included consistently. The resulting corrections are close to unity with typical
uncertainties between 5 and 10%.
Measurements of misidentification rates can be carried out by selecting a dijet
sample with two high- p T jets, which are dominated by gluon and light-flavour jets.
In order to test different flavour compositions, Z /γ +jets selections are used as well.
Due to the high p T threshold of unprescaled jet triggers, measurements on dijet
3 Jet Substructure at the LHC
approximately 20%. There exist other approaches to cover the transition from low
to high Lorentz boosts, using a single algorithm. In the HTTv2 algorithm, the jet
size is reduced until an optimal size R opt is found, defined by the fractional jet mass
contained in the smaller jet. This results in better performance at high p T , while keeping a low misidentification rate at low p T . Similar performance is obtained with the
HOTVR algorithm, although with less algorithmic complexity. A comparison of the
efficiencies and misidentification rates for the CMSTT, HTT, HTTv2 and HOTVR
as a function of p T is shown in Fig. 3.11 (right). This comparison has been obtained
with the parameter settings described in [245]. The tagger working points have been
chosen to obtain an efficiency of 30% at p T = 700 GeV. The turn-on of the HOTVR
algorithm is not as sharp as for the HTT or HTTv2 taggers, but the misidentification
rate is smaller in this region. Overall, the HOTVR tagger achieves a misidentification
rate comparable to the one of the CMSTT over the full range in p T , while achieving
similar efficiencies as the HTT or HTTv2 taggers at low p T . The high efficiency
of the HOTVR tagger starting from p T > 300 GeV, together with a flat background
efficiency, will make this tagger an important tool in analyses targeting intermediate
to high top quark boosts.
An important step towards the commissioning of top taggers within an experiment
are measurements of the efficiency and misidentification rate in real collision data.
Generally, high-purity samples of top jets in data are obtained using a tight signal
selection (a lepton, well-separated from a high- p T large-R jet, and an additional btagged jet) to ensure that events contain a fully-merged top quark decay in a single
large-R jet. This can never be fully achieved, as no requirements on the substructure of the large-R jet can be imposed without biasing the efficiency measurement.
This results in an efficiency measurement that will be based on a sample containing
partially-merged and non-merged top quark decays. These can be subtracted from
the efficiency measurement by using simulated events, as done in a measurement
by ATLAS [520]. This has the drawback of relying on a specific simulation and
the ambiguous definition of a fully merged top quark decay. By not correcting for
non-merged top quark decays, efficiency values are obtained smaller than the ones
suggested by ROC curve studies, see for example [459]. Instead of subtracting the
t backgrounds, recent measurements perform a simultaneous extraction of the efficiencies for fully and partially merged categories [525, 526, 530, 531]. This is done
by separating events into a passing and a failing region, where pass and fail are determined by a selection on τ 32 . The obtained distributions in the (groomed) jet mass are
used for a template fit of the fully merged, semi-merged and un-merged categories
to data, as shown in Fig. 3.12. The jet mass is used since the three contributions
have different shapes in this distribution, allowing for a precise determination of the
top tagging efficiencies. An additional advantage is that systematic uncertainties can
be included consistently. The resulting corrections are close to unity with typical
uncertainties between 5 and 10%.
Measurements of misidentification rates can be carried out by selecting a dijet
sample with two high- p T jets, which are dominated by gluon and light-flavour jets.
In order to test different flavour compositions, Z /γ +jets selections are used as well.
Due to the high p T threshold of unprescaled jet triggers, measurements on dijet
