submit data files along with the text of a publication. However, since there is as yet
no standard and little infrastructure adopted by mainstream publishers for dealing
with this data, most of it remains in inaccessible files (e.g. PDF documents). An
appropriate answer to this dilemma is the semantic web.
In conjunction with the data, the semantic web includes a vocabulary for
describing the data. This is where semantics comes to the fore. This vocabulary for
a field such as Quantum Chemistry needs to be encoded in a formal language. The
Web Ontology Language (OWL) [3] is such a language and the “vocabulary” is
really an ontology describing the formal language of a specific domain such as
computational or quantum chemistry. Chemical Semantics has created such an
ontology which is referred to as the Gainesville Core (http://purl.org/chem/gc). This
first ontology for quantum chemistry will need to be modified and extended by
scientists in the field as basic ontology ideas become more prevalent in chemistry.
Our ontology is meant to be placed in the public domain so that the quantum
chemistry community can both contribute to it and use it. Developing the Gainesville Core is an ongoing project with release numbers.
The semantic web allows scientists to publish their data in a structured way such
that it can be found and used by anyone with access to the World Wide Web.
Chemical Semantics, Inc. is creating client and server software that allows scientists
to automate their publishing of data into a modern graph database [4], i.e. a Giant
Global Graph (GGG), so that they and others can share their data and use it in a way
that is not now possible with the existing World Wide Web (WWW). Adding
semantics to scientific data and making that data available on the semantic web
makes it possible to do science in a new collaborative way that has the potential to
change science forever [1].
1.1 Publishing Quantum Chemistry Data
To make the technology of the semantic web available to scientists, one has to
“publish data” in a way that is related to the standard model of journal publishing.
That is, one puts the data (not journal text) into an appropriate form, sends it to a
publisher (of data not journal text), waits for its publication, and then informs
colleagues that they can access the data (not journal article) in a standard fashion
(more and more via the web rather than via hard copy).
The appropriate form for data described above has been clearly defined by the
World Wide Web Consortium (W3C) [5] as the Resource Description Framework
(RDF) standard [6]. This standard is a Graph Database where RDF statements all
take the “triple” form (subject, predicate, and object). For example, (water,
has_boiling_point, 100) is of this triple form. That is, there is a graph arc called
“has_boiling_point” which points from a “water” subject node to a “100” object
node. This form surpasses the normal relational database in its applicability to the
web. Any and all scientific data can be put into this form and data on the semantic
web is stored in servers referred to as triple stores. These triple stores may contain
billions of triples.
4
B. Wang et al.
no standard and little infrastructure adopted by mainstream publishers for dealing
with this data, most of it remains in inaccessible files (e.g. PDF documents). An
appropriate answer to this dilemma is the semantic web.
In conjunction with the data, the semantic web includes a vocabulary for
describing the data. This is where semantics comes to the fore. This vocabulary for
a field such as Quantum Chemistry needs to be encoded in a formal language. The
Web Ontology Language (OWL) [3] is such a language and the “vocabulary” is
really an ontology describing the formal language of a specific domain such as
computational or quantum chemistry. Chemical Semantics has created such an
ontology which is referred to as the Gainesville Core (http://purl.org/chem/gc). This
first ontology for quantum chemistry will need to be modified and extended by
scientists in the field as basic ontology ideas become more prevalent in chemistry.
Our ontology is meant to be placed in the public domain so that the quantum
chemistry community can both contribute to it and use it. Developing the Gainesville Core is an ongoing project with release numbers.
The semantic web allows scientists to publish their data in a structured way such
that it can be found and used by anyone with access to the World Wide Web.
Chemical Semantics, Inc. is creating client and server software that allows scientists
to automate their publishing of data into a modern graph database [4], i.e. a Giant
Global Graph (GGG), so that they and others can share their data and use it in a way
that is not now possible with the existing World Wide Web (WWW). Adding
semantics to scientific data and making that data available on the semantic web
makes it possible to do science in a new collaborative way that has the potential to
change science forever [1].
1.1 Publishing Quantum Chemistry Data
To make the technology of the semantic web available to scientists, one has to
“publish data” in a way that is related to the standard model of journal publishing.
That is, one puts the data (not journal text) into an appropriate form, sends it to a
publisher (of data not journal text), waits for its publication, and then informs
colleagues that they can access the data (not journal article) in a standard fashion
(more and more via the web rather than via hard copy).
The appropriate form for data described above has been clearly defined by the
World Wide Web Consortium (W3C) [5] as the Resource Description Framework
(RDF) standard [6]. This standard is a Graph Database where RDF statements all
take the “triple” form (subject, predicate, and object). For example, (water,
has_boiling_point, 100) is of this triple form. That is, there is a graph arc called
“has_boiling_point” which points from a “water” subject node to a “100” object
node. This form surpasses the normal relational database in its applicability to the
web. Any and all scientific data can be put into this form and data on the semantic
web is stored in servers referred to as triple stores. These triple stores may contain
billions of triples.
4
B. Wang et al.
