package ought perhaps to have such a button! The founders of Chemical Semantics,
Inc. have a historical tie to Hypercube, Inc. and as such have used HyperChem to
illustrate the basic ideas. What this button in a “CompChem” package does is create
an XML file structure that includes the information about the current molecular
system and any current calculation results resident in the package and sends the data
to the Chemical Semantics, Inc. portal where it is published, given the Authors,
Title, Abstract, Login information and other details that are part of the global setup
prior to hitting the “Publish” button. Our current XML file structure is called CSX
and is related in historical terms to the Chemical Markup Language (CML) file
structure [12]. We use CSX because CML currently does not have certain properties
that we consider not only desirable, but mandatory, such as the ability to deal with
residues as independent fundamental units of biological molecules. Also, CML
does not define many of the quantum chemistry concepts that we consider critical,
We believe CSX encompasses CML as a subset and we could generate a CML file
from CSX easily.
Our Chemical Semantics portal accepts data using a REST or SOAP [13] protocol and then publishes the data on its servers. The data is available to anyone
around the world with an account at the portal. The details of the portal are
described in another section below.
1.2 Searching the Data
Our portal allows users to access data at the portal based upon various rules, search
criteria, etc. Semantic web data is usually searched for using a SPARQL Protocol
and RDF Query Language (SPARQL) [14]. Note the recursive acronym.
A SPARQL end point is maintained at the portal which allows queries of various
kinds. SPARQL has features in common with the SQL query language and is
relatively easy to use but requires some experience in forming queries. A natural
language front end would be desirable and many groups are involved in developing
such front ends. A query, for example, could ask how many Density Functional
Theory (DFT) calculations have been done on a specific molecule and with which
functionals and what final total energies. As opposed to querying relational database silos, the query could potentially survey the whole world’s set of such calculations and return with a table of these.
Because the data stored by Chemical Semantics includes a relatively unlimited
number of triples, searching can be very exhaustive while still being very focused.
The data is defined by the ontology so that a search does not return irrelevant
results. Because the data is held by a graph database where a graph node (resource)
uses an arrow to point to another resource, the arrow can simply point to a resource
in a second graph from the first graph and unlike a relational database the data from
different graphs can be merged trivially. This “federation” allows searches to really
use a Giant Global Graph (GGG). As part of the publish activity, data on our portal
can be tagged as private, protected, or public. Any search obviously can return
6
B. Wang et al.
Précédent

- 18/406

Suivant