206
F. Firouzi and B. Farahani
of Elasticsearch can be started to form a cluster with any number of desired nodes.
Each node within the cluster has knowledge of the other nodes within the same
cluster because they communicate directly with one another via TCP. This is referred
to as fully connected mesh topology. Every node within the cluster handles one or
more roles, namely, a data node, a master node, a client node, or an ingest node [24]:
• Master Node – Responsible for constructing or eliminating indices and adding
or removing nodes. Every time the cluster state changes, the master node notifies
the other nodes within the cluster about the change. Each cluster contains only
one master node at a time.
• Data Node – Each node contains a data shard (a partition of whole data) and
handles data-related operations including creating, reading, deleting, updating,
searching, and aggregating. A cluster may contain many data nodes. If one
data node within the cluster should stop, the cluster continues operations and
rearranges that node’s data on the additional nodes.
• Client Node – It handles sending cluster-related requests to the master node as
well as sending data-related requests to data nodes by serving as a “smart router.”
A client node does not contain data and is unable to become a master node.
• Ingest Node – Before actual indexing occurs, it handles preliminary processing
of documents.
Index An index is a collection of documents with similar characteristics. An index
is represented by a unique name which is used to refer to the index while executing
operations such as index searches, updates, and deletions. A cluster may contain as
many indexes as desired. In Elasticsearch, the index is comparable to the database
schema in RDBMS and can be thought of as a set of tables with logical organization
or grouping. In the same way, Type can be considered equivalent to Table and
Document as equivalent to Row in RDBMS.
Document A document is simply a collection of fields organized in a specific
JSON format. Each document belongs to a type, and it is stored in an index,
associated with a unique identifier (UID).
Type/Mapping Type or mapping refers to a collection of documents that share a
common set of fields found in the same index. For example, an index containing
social networking application data can contain specific user profile data, another
document containing messaging data, and yet another for social media comment
data. We should note that Elasticsearch recently indicated that it would no longer
be possible to include multiple types in an index, with the concept of types being
eliminated in a later version.
Shards and Replicas An index is able to store massive amounts of data exceeding
the hardware capabilities of the node. For example, an index containing 1 billion
documents and requiring 1 TB of disk space may not fit on the node disk or become
too slow to serve search requests from a node. To be able to address this large
amount of data, Elasticsearch gives users the ability to horizontally subdivide an
index into multiple, smaller horizontal pieces known as shards. Because of this,
Précédent

- 213/647

Suivant