82
4 Google as God?
4.12.4 Ambiguity
Information can have different meanings. In many cases, the correct interpretation
can be found only with some additional piece of information: the context. This
contextualization is often difficult and not always available when needed. Different
pieces of information may also be inconsistent, without any chance to resolve the
conflict.
Another typical problem of “data mining” is that data might be plentiful, but
nevertheless incomplete and non-representative. Moreover, a lot of data might be
wrong, because of measurement errors, data manipulation, misinterpretations, or the
application of unsuitable procedures.
4.12.5 Information Overload
Having a lot of data does not necessarily mean that we’ll see the world more accurately. A typical problem is known as “over-fitting”, where a model with many
parameters is fitted to the fine details of a data set in ways that are actually not meaningful. In such a case, a model with less parameters might provide better predictions.
Spurious correlations are a somewhat similar problem: we tend to see patterns that
actually don’t have a meaning.
58
It seems that we are currently moving from a situation where we had too little
data about the world to a situation where we have too much. It’s like moving from
darkness, where we can’t see enough, to a world flooded with light, in which we are
blinded. We will, therefore, need “digital sunglasses”: information filters that will
extract the relevant information for us. But as the gap between the data that exists and
the data we can analyze increases, it might become harder to pay attention to those
things that really matter. Although computer power increases exponentially, we will
only be able to process an ever-decreasing fraction of data, since the storage capacity
grows even faster. In other words, there will be increasing volumes of data that will
never be touched (which I call “dark data”). To reap the benefits of the information
age, we must reduce information biases and pollution, otherwise we will increasingly
make mistakes.
4.12.6 Herding
In many situations, which are not entirely well defined, people tend to follow the
decisions and actions of others. This produces undesirable herding effects. The Nobel
laureates in economics, George Akerlof (*1940) and Robert Shiller (*1946), have
called this problematic behavior “animal spirits”, but the idea of herding goes back
58 See http://www.tylervigen.com/spurious-correlations for some examples.
4 Google as God?
4.12.4 Ambiguity
Information can have different meanings. In many cases, the correct interpretation
can be found only with some additional piece of information: the context. This
contextualization is often difficult and not always available when needed. Different
pieces of information may also be inconsistent, without any chance to resolve the
conflict.
Another typical problem of “data mining” is that data might be plentiful, but
nevertheless incomplete and non-representative. Moreover, a lot of data might be
wrong, because of measurement errors, data manipulation, misinterpretations, or the
application of unsuitable procedures.
4.12.5 Information Overload
Having a lot of data does not necessarily mean that we’ll see the world more accurately. A typical problem is known as “over-fitting”, where a model with many
parameters is fitted to the fine details of a data set in ways that are actually not meaningful. In such a case, a model with less parameters might provide better predictions.
Spurious correlations are a somewhat similar problem: we tend to see patterns that
actually don’t have a meaning.
58
It seems that we are currently moving from a situation where we had too little
data about the world to a situation where we have too much. It’s like moving from
darkness, where we can’t see enough, to a world flooded with light, in which we are
blinded. We will, therefore, need “digital sunglasses”: information filters that will
extract the relevant information for us. But as the gap between the data that exists and
the data we can analyze increases, it might become harder to pay attention to those
things that really matter. Although computer power increases exponentially, we will
only be able to process an ever-decreasing fraction of data, since the storage capacity
grows even faster. In other words, there will be increasing volumes of data that will
never be touched (which I call “dark data”). To reap the benefits of the information
age, we must reduce information biases and pollution, otherwise we will increasingly
make mistakes.
4.12.6 Herding
In many situations, which are not entirely well defined, people tend to follow the
decisions and actions of others. This produces undesirable herding effects. The Nobel
laureates in economics, George Akerlof (*1940) and Robert Shiller (*1946), have
called this problematic behavior “animal spirits”, but the idea of herding goes back
58 See http://www.tylervigen.com/spurious-correlations for some examples.
