Turing’s Test 109
it more efficiently than any human could, and also defeating a human
professional Go player in 2015. Unlike chess, which had demonstrated
itself capable of exercising iterative heuristics – being able to process an
increasing number of moves ahead – Go was believed not to be amenable
to such a process. With regard to reading, however, the best technique
turned out to be feeding huge data sets of information to the algorithm,
with experts at DeepMind feeding large quantities of Daily Mail and
CNN articles to learn to read. To demonstrate that comprehension was
possible, Karl Moritz Hermann and a team of engineers working at
Google set up the system to allow it to extract bullet points from text
without simply repeating sentences within the data set.
44
With regard to such machine comprehension, in early 2018, teams
from Microsoft and Alibaba claimed independently that they had created AIs that could match human performance on the Stanford Question Answering Dataset (SQuAD). SQuAD, currently at iteration 2.0,
comprises 150,000 questions which require the reader to comprehend
a corresponding passage of text before they can be answered. To make
the task more difficult, 50,000 of those questions are deliberately unanswerable to ensure that human and machine readers are clear about
what they don’t know as well as what they do. Microsoft, in a blog post
in January 2018, announced that it had achieved a SQuAD score of
82.6 per cent (comparable to 82.3 per cent for humans) and that it was
jointly tied with Alibaba.
45
This post led to a slew of headlines about
machines replacing humans, with Newsweek estimating “millions of
jobs at risk”,
46
but more critical commentators noted that the SQuAD
scoring system relied on Mechanical Turk workers paid $9 an hour to
answer questions, who would probably be less motivated to find correct answers than machine systems,
47
and that while the data set looks
challenging (with questions on Reformation theology or the concept of
civil disobedience), in practice answers rely not on any knowledge of
the subject but instead being able to match patterns. For example, as
James Vincent remarked in The Verge, while a question such as “Whose
authority does Luther’s theology oppose?” may seem tough, the fact that
a reading passage includes the sentence “[Luther’s] theology challenged
the authority and office of the Pope” makes it clear that this entire test
operates around a restricted notion of comprehension.
48
As an expert in
NLP, Yoav Goldberg, told Vincent, the test was designed as “a benchmark for machine learning methods”, not as for comparisons to human
readers. Researchers believed in the 1950s that automated translation
was just around the corner: today, we can see some very effective results from machine translation, but NLP and NLG generally work better
within clearly defined parameters rather than entirely unsupervised on
completely open texts. Likewise, there have been considerable advances
in automated or algorithmic journalism, in the past half-decade – but
the successes are clearest when dealing with story generation in limited
Précédent

- 119/200

Suivant