of the content of a Web page instead of the URL at the time of request by a user. China may
be the sole exception to this rule. Our research has documented an elaborate network of controls including keyword-based URL filtering; it may be the case that Chinese filtering systems
can be triggered based upon the presence of keywords within a Web page’s content. OpenNet Initiative researcher Steven Murdoch along with his colleagues Richard Clayton and Robert N. M. Watson have published a paper that describes in detail the workings of the ‘‘Great
Firewall of China,’’ including this dynamic filtering based on Web page content. As Clayton,
Murdoch, and Watson note, ‘‘We have demonstrated that the ‘Great Firewall of China’ relies
on inspecting packets for specific content.’’
16 Yemen and Iran also have the capacity to block
sites based on keywords in URLs. Commercial software packages such as SmartFilter make
such URL filtering trivial. As a general rule, with the possible exception of China, access to a
site is based on its URL; if a URL has triggered a block, one could take down the offensive
content within the page and replace it with the most innocuous material possible, and the
original link will continue to trigger a block.
URL-based blocking does not, however, require the identification of every page that is to
be blocked. Our research indicates that the most prevalent form of blocking is at the domain
level. Once a state has identified www.playboy.com as undesirable, the logical step is to deny
all requests to that domain, whether http://playboy.com/playmates/2003/may.html or playboy
.com/articles/interviews/index.html.
The parallel between the URL-based approach with the approach of the traditional censor is
that the domain is deemed on the whole undesirable, and the censor makes no effort to disaggregate the content within. The decision is most complicated when a single domain hosts
truly disparate content, such as free hosting sites like Geocities or Angelfire, blogging domains like Blogspot or Blogger, community sites like Google or Yahoo! groups, or university
sites like mit.edu that can include student home pages about subjects like Tibet. Within these
realms, our research found ample evidence of both blocking of the entire domain and
selected blocking of subsites, or pages within the domain. Such blocking is discussed in the
respective state reports in the appendix to this book. The Berkman Center’s Web site at http://
cyber.law.harvard.edu was blocked in China in 2002 after our first report of Internet filtering
was placed there. (The powers that be in Harvard’s central university administration declined
to repost the study at www.harvard.edu.)
We have also observed several other means of URL-based filtering. As the results presented
in chapter 1 show, several states—including China, Iran, Myanmar, and Yemen—block access
to all URLs containing particular strings of letters (such as ‘‘ass’’), whether such banned terms
appear in the domain or in superfluous characters at its end. Those sites’ IP addresses are
independently blocked, as blocked domains could otherwise be accessed via this method.
Some blocking approaches are cruder still. We observed that the United Arab Emirates and
Syria blocked every site found within the Israeli top-level country code domain: no pages from
any domain ending in ‘‘.il’’ were accessible there.
17
Internet Filtering: The Politics and Mechanisms of Control
37
be the sole exception to this rule. Our research has documented an elaborate network of controls including keyword-based URL filtering; it may be the case that Chinese filtering systems
can be triggered based upon the presence of keywords within a Web page’s content. OpenNet Initiative researcher Steven Murdoch along with his colleagues Richard Clayton and Robert N. M. Watson have published a paper that describes in detail the workings of the ‘‘Great
Firewall of China,’’ including this dynamic filtering based on Web page content. As Clayton,
Murdoch, and Watson note, ‘‘We have demonstrated that the ‘Great Firewall of China’ relies
on inspecting packets for specific content.’’
16 Yemen and Iran also have the capacity to block
sites based on keywords in URLs. Commercial software packages such as SmartFilter make
such URL filtering trivial. As a general rule, with the possible exception of China, access to a
site is based on its URL; if a URL has triggered a block, one could take down the offensive
content within the page and replace it with the most innocuous material possible, and the
original link will continue to trigger a block.
URL-based blocking does not, however, require the identification of every page that is to
be blocked. Our research indicates that the most prevalent form of blocking is at the domain
level. Once a state has identified www.playboy.com as undesirable, the logical step is to deny
all requests to that domain, whether http://playboy.com/playmates/2003/may.html or playboy
.com/articles/interviews/index.html.
The parallel between the URL-based approach with the approach of the traditional censor is
that the domain is deemed on the whole undesirable, and the censor makes no effort to disaggregate the content within. The decision is most complicated when a single domain hosts
truly disparate content, such as free hosting sites like Geocities or Angelfire, blogging domains like Blogspot or Blogger, community sites like Google or Yahoo! groups, or university
sites like mit.edu that can include student home pages about subjects like Tibet. Within these
realms, our research found ample evidence of both blocking of the entire domain and
selected blocking of subsites, or pages within the domain. Such blocking is discussed in the
respective state reports in the appendix to this book. The Berkman Center’s Web site at http://
cyber.law.harvard.edu was blocked in China in 2002 after our first report of Internet filtering
was placed there. (The powers that be in Harvard’s central university administration declined
to repost the study at www.harvard.edu.)
We have also observed several other means of URL-based filtering. As the results presented
in chapter 1 show, several states—including China, Iran, Myanmar, and Yemen—block access
to all URLs containing particular strings of letters (such as ‘‘ass’’), whether such banned terms
appear in the domain or in superfluous characters at its end. Those sites’ IP addresses are
independently blocked, as blocked domains could otherwise be accessed via this method.
Some blocking approaches are cruder still. We observed that the United Arab Emirates and
Syria blocked every site found within the Israeli top-level country code domain: no pages from
any domain ending in ‘‘.il’’ were accessible there.
17
Internet Filtering: The Politics and Mechanisms of Control
37
