9 Towards the “Shape” of Cosmological Observables and the String …
227
fact, one can make contact to filtrations defined via point sets X ⊂ R
m by considering
the distance function
f X (y) = inf
x∈X
x − y
(9.4)
Now letting P D n ( f ) represent the persistence diagram corresponding to the ndimensional persistent homology of the sublevel filtration of f , one can show that
d B (P D n ( f ), P D n (g)) ≤ f − g ∞
(9.5)
where d B is the so-called bottleneck distance between persistence diagrams, defined
by
d B (P D n ( f ), P D n (g)) = inf
η
sup
x∈P D n ( f )
x − η(x)| ∞
(9.6)
where η is a bijection between P D n ( f ) and P D n (g), and the finite number of nonzero
persistence cycles in a given persistence diagram is supplemented by an infinite
number of zero-persistence cycles along the “birth=death” line of the persistence
diagram. In other words, if the input data of a filtration is perturbed, the change
in the PD is provably small. Intuitively, this is because small noise will not create
or destroy significant topological features. Thus we see how the introduction of
persistence leads to significant robustness.
This guarantee of robustness might be compared with the problem of adversarial
perturbations for machine learning frameworks [97]. Here, imperceptible noise can
cause neural networks and other ML architectures to e.g.. misclassify images in
dramatic ways. One might hope that a machine learning pipeline utilizing persistent
homology can be made more robust to such perturbations. Note also that slicings of
PDs like the persistent Betti numbers do not inherit stability from PDs.
9.2.5 Coupling to Machine Learning, Statistical Inference
As scatter plots, PDs do not immediately couple nicely to machine learning pipelines.
However, there are several routes one can take to define nicer mathematical objects.
One derived statistic that does inherit stability is the persistence image [4], which
can be used for statistical learning. Roughly speaking, the persistence image is a
persistence diagram where individual cycles have been smoothed via some specified
kernel and the resulting scalar function has been binned as a 2-dimensional histogram.
To inherit the stability of PDs, one also multiplies the kernel of a given cycle by a
factor that depends on the cycle’s persistence, for example log(1 + ν death − ν birth ).
The resulting object lives in a vector space and as such naturally couples to statistical
learning pipelines. Another persistence-based object suitable for statistical inference
227
fact, one can make contact to filtrations defined via point sets X ⊂ R
m by considering
the distance function
f X (y) = inf
x∈X
x − y
(9.4)
Now letting P D n ( f ) represent the persistence diagram corresponding to the ndimensional persistent homology of the sublevel filtration of f , one can show that
d B (P D n ( f ), P D n (g)) ≤ f − g ∞
(9.5)
where d B is the so-called bottleneck distance between persistence diagrams, defined
by
d B (P D n ( f ), P D n (g)) = inf
η
sup
x∈P D n ( f )
x − η(x)| ∞
(9.6)
where η is a bijection between P D n ( f ) and P D n (g), and the finite number of nonzero
persistence cycles in a given persistence diagram is supplemented by an infinite
number of zero-persistence cycles along the “birth=death” line of the persistence
diagram. In other words, if the input data of a filtration is perturbed, the change
in the PD is provably small. Intuitively, this is because small noise will not create
or destroy significant topological features. Thus we see how the introduction of
persistence leads to significant robustness.
This guarantee of robustness might be compared with the problem of adversarial
perturbations for machine learning frameworks [97]. Here, imperceptible noise can
cause neural networks and other ML architectures to e.g.. misclassify images in
dramatic ways. One might hope that a machine learning pipeline utilizing persistent
homology can be made more robust to such perturbations. Note also that slicings of
PDs like the persistent Betti numbers do not inherit stability from PDs.
9.2.5 Coupling to Machine Learning, Statistical Inference
As scatter plots, PDs do not immediately couple nicely to machine learning pipelines.
However, there are several routes one can take to define nicer mathematical objects.
One derived statistic that does inherit stability is the persistence image [4], which
can be used for statistical learning. Roughly speaking, the persistence image is a
persistence diagram where individual cycles have been smoothed via some specified
kernel and the resulting scalar function has been binned as a 2-dimensional histogram.
To inherit the stability of PDs, one also multiplies the kernel of a given cycle by a
factor that depends on the cycle’s persistence, for example log(1 + ν death − ν birth ).
The resulting object lives in a vector space and as such naturally couples to statistical
learning pipelines. Another persistence-based object suitable for statistical inference
