Chapter 11
Summary
The way in which the quality of a model depends on the amount of data that it is
built from is actually quite subtle. For example, if we want to build a prediction
boundary between data of two classes, then having more examples of each class is
a good thing, but the accuracy of the boundary does not increase linearly with the
number of examples. Rather, it increases quickly and then flattens. The perspective
of support vector machines shows why — new records are only useful to improve
the boundary when they are support vectors; as the data increases, each new record
becomes less and less likely to be a support vector.
If we want to find a clustering, then more data will tend to produce a better
clustering, but without some data, it is impossible even to tell whether a few scattered
records are a cluster or not. So while a few records from each class can already locate
the prediction boundary between the classes quite well, a few records do not enable
any conclusions for a clustering. We could say that the clusters are emergent from
the data in this weak sense.
When it comes to social networks, the requirement for lots of data becomes
even more important. Knowing a few scattered relationships does not enable any
general conclusions. As the number of known relationships increases, brand new
properties of the social network begin to appear. As the network of relationships
becomes connected, we can define the diameter of the network. Once the network
is more or less complete, we can compute the degrees of the nodes and discover that
they tend to follow a power law. We can differentiate those nodes that are somehow
central and those nodes that are peripheral.
Because social networks are generated by humans, making individual, local
decisions about whom to relate to, and what form and intensity that relationship will
take, these emergent properties are particularly revealing and important, because they
provide insights at a more abstract level about why these relationships are formed.
From the perspective of each individual, the decision to form a relationship is a
local and private one. The emergent structures of social networks show that this
157
Summary
The way in which the quality of a model depends on the amount of data that it is
built from is actually quite subtle. For example, if we want to build a prediction
boundary between data of two classes, then having more examples of each class is
a good thing, but the accuracy of the boundary does not increase linearly with the
number of examples. Rather, it increases quickly and then flattens. The perspective
of support vector machines shows why — new records are only useful to improve
the boundary when they are support vectors; as the data increases, each new record
becomes less and less likely to be a support vector.
If we want to find a clustering, then more data will tend to produce a better
clustering, but without some data, it is impossible even to tell whether a few scattered
records are a cluster or not. So while a few records from each class can already locate
the prediction boundary between the classes quite well, a few records do not enable
any conclusions for a clustering. We could say that the clusters are emergent from
the data in this weak sense.
When it comes to social networks, the requirement for lots of data becomes
even more important. Knowing a few scattered relationships does not enable any
general conclusions. As the number of known relationships increases, brand new
properties of the social network begin to appear. As the network of relationships
becomes connected, we can define the diameter of the network. Once the network
is more or less complete, we can compute the degrees of the nodes and discover that
they tend to follow a power law. We can differentiate those nodes that are somehow
central and those nodes that are peripheral.
Because social networks are generated by humans, making individual, local
decisions about whom to relate to, and what form and intensity that relationship will
take, these emergent properties are particularly revealing and important, because they
provide insights at a more abstract level about why these relationships are formed.
From the perspective of each individual, the decision to form a relationship is a
local and private one. The emergent structures of social networks show that this
157
