228
Barbara Carminati and Elena Ferrari
to each node. The publisher is thus able to perform (approximate) queries directly
on the encrypted documents, by exploiting the partitioning ids. Information on partition ids are contained in the SE-ENC document. To adapt this scheme to XML,
we have, first of all, to deal with partitioning techniques for textual data, which
play a key role in XML (e.g. element data content). Our approach to encrypt textual data contained in an XML document is based on two main steps. The first is
keyword extraction and it is performed mainly to reduce the search space. Then,
each keyword is considered as a distinct partition, whose id is obtained by applying
a one-way hash function to the keyword itself. Properties of hash functions ensure
that it is almost impossible to find two different keywords with the same id, and
that it is computationally infeasible for the publisher to obtain the original keyword,
given the corresponding id. The only drawback of this scheme is that the publisher
could infer information on the clear- text by analyzing the distribution of hash values.
Therefore, in [9] we have extended this method to be robust against data dictionary
attacks.
To submit queries against encrypted data, a user must translate the query into
an encrypted form which is understandable by publishers. This is done by the query
translator module (see [7] for more details) that receives as input a user query and
returns as output one or more ciphered queries to be submitted to publishers. Main
steps done by the query translator are the completion of the user query (since the user
may be allowed to see only a view of the requested document), the encryption of the
tag and attribute names contained in the input query, and the translation of the query
predicates in terms of partition ids. Information about partitioning functions and ids
generation mechanisms is stored in the user entry of the owner’s directory server.
As a final remark, note that the fact that a user submits encrypted queries to
publishers ensures a certain degree of user’s privacy in that publishers do not know
the details of the submitted queries.
10.4.3 Completeness
Completeness verification is based on an XML-data structure called query template,
which is generated by the owner for each outsourced document and downloaded by
clients from the owner’s directory the first time they query the corresponding document. The query template is generated by the owner, by applying a simple XSLT
transformation [28] on the corresponding SE-ENC document, which prunes from the
SE-ENC document the encrypted data contents and security-related information not
necessary for completeness verification. The query template is signed by owners with
a Merkle signature. Since the query template is generated from the SE-ENC document, it contains the structure of the corresponding document selectively encrypted
according to the owner’s access control policies, plus additional information needed
for completeness verification, that is, partition ids and access control information.
The idea is that by exploiting the same query processing strategy used by publishers (i.e. based on partition ids), the user is able to perform on the query template
Barbara Carminati and Elena Ferrari
to each node. The publisher is thus able to perform (approximate) queries directly
on the encrypted documents, by exploiting the partitioning ids. Information on partition ids are contained in the SE-ENC document. To adapt this scheme to XML,
we have, first of all, to deal with partitioning techniques for textual data, which
play a key role in XML (e.g. element data content). Our approach to encrypt textual data contained in an XML document is based on two main steps. The first is
keyword extraction and it is performed mainly to reduce the search space. Then,
each keyword is considered as a distinct partition, whose id is obtained by applying
a one-way hash function to the keyword itself. Properties of hash functions ensure
that it is almost impossible to find two different keywords with the same id, and
that it is computationally infeasible for the publisher to obtain the original keyword,
given the corresponding id. The only drawback of this scheme is that the publisher
could infer information on the clear- text by analyzing the distribution of hash values.
Therefore, in [9] we have extended this method to be robust against data dictionary
attacks.
To submit queries against encrypted data, a user must translate the query into
an encrypted form which is understandable by publishers. This is done by the query
translator module (see [7] for more details) that receives as input a user query and
returns as output one or more ciphered queries to be submitted to publishers. Main
steps done by the query translator are the completion of the user query (since the user
may be allowed to see only a view of the requested document), the encryption of the
tag and attribute names contained in the input query, and the translation of the query
predicates in terms of partition ids. Information about partitioning functions and ids
generation mechanisms is stored in the user entry of the owner’s directory server.
As a final remark, note that the fact that a user submits encrypted queries to
publishers ensures a certain degree of user’s privacy in that publishers do not know
the details of the submitted queries.
10.4.3 Completeness
Completeness verification is based on an XML-data structure called query template,
which is generated by the owner for each outsourced document and downloaded by
clients from the owner’s directory the first time they query the corresponding document. The query template is generated by the owner, by applying a simple XSLT
transformation [28] on the corresponding SE-ENC document, which prunes from the
SE-ENC document the encrypted data contents and security-related information not
necessary for completeness verification. The query template is signed by owners with
a Merkle signature. Since the query template is generated from the SE-ENC document, it contains the structure of the corresponding document selectively encrypted
according to the owner’s access control policies, plus additional information needed
for completeness verification, that is, partition ids and access control information.
The idea is that by exploiting the same query processing strategy used by publishers (i.e. based on partition ids), the user is able to perform on the query template
