4 Architecting IoT Cloud
205
• Measurement – This can also be thought of as an SQL table because the
measurement serves as a container for the time column, fields, and tags.
• Retention Policy – Retention policy describes the amount of time that an
InfluxDB database stores data (duration) and how many data copies are kept
within the cluster (replication factor). Data that is older than the given duration
is automatically removed from the database. Generally, the minimum retention
duration is 1 hour and the maximum duration is infinite.
• Sharding – Sharding refers to the horizontal partitioning of data with each
partition known as a shard. InfluxDB also utilizes sharding to address the
scalability problem. Each shard holds compressed, actual data of a particular
series set. A shard can belong to only one shard group, and there may be many
shards in a particular group. All points within a set series in a shard group will
be stored in the same shard. The duration parameter specifies the amount of time
a shard group covers. In other words, the shard duration determines the specific
interval. For example, if shard duration is set to 1 week, each shard group will
cover 1 week and include all data points with timestamps in that week timeframe.
4.6.1.5 Elasticsearch
Elasticsearch (ES) is a highly popular soft real-time search engine used by largescale organizations including The Guardian, GitHub, StackOverflow, and Wikipedia.
ES is considered as a document-oriented database created to hold, manage, and
retrieve semi-structured or document-oriented data. Table 4.5 compares Elasticsearch with RDBMS [17, 22, 23].
There are a few basic concepts that one must grasp in order to understand the
functionality and structure of Elasticsearch better.
Cluster Clusters are a collection of servers (nodes) that, when networked together,
hold the complete data set and provide federated indexing and search capability for
all connected servers.
Node A node is a solitary server that contains a portion of data and can contribute
to the cluster’s querying and indexing tasks. Note that each cluster is identified with
a name and a node can be assigned to a particular cluster based on the cluster’s
name. Starting an instance of Elasticsearch from scratch results in the creation of a
cluster with only one node. When the second instance of Elasticsearch starts (with
the same “cluster name”), the cluster will contain two nodes. Additional instances
Table 4.5 Relationship of
RDBMS terminology with
Elasticsearch
RDBMS Elasticsearch
Database Index
Table
Mapping
Field
Field
Row
JSON object
205
• Measurement – This can also be thought of as an SQL table because the
measurement serves as a container for the time column, fields, and tags.
• Retention Policy – Retention policy describes the amount of time that an
InfluxDB database stores data (duration) and how many data copies are kept
within the cluster (replication factor). Data that is older than the given duration
is automatically removed from the database. Generally, the minimum retention
duration is 1 hour and the maximum duration is infinite.
• Sharding – Sharding refers to the horizontal partitioning of data with each
partition known as a shard. InfluxDB also utilizes sharding to address the
scalability problem. Each shard holds compressed, actual data of a particular
series set. A shard can belong to only one shard group, and there may be many
shards in a particular group. All points within a set series in a shard group will
be stored in the same shard. The duration parameter specifies the amount of time
a shard group covers. In other words, the shard duration determines the specific
interval. For example, if shard duration is set to 1 week, each shard group will
cover 1 week and include all data points with timestamps in that week timeframe.
4.6.1.5 Elasticsearch
Elasticsearch (ES) is a highly popular soft real-time search engine used by largescale organizations including The Guardian, GitHub, StackOverflow, and Wikipedia.
ES is considered as a document-oriented database created to hold, manage, and
retrieve semi-structured or document-oriented data. Table 4.5 compares Elasticsearch with RDBMS [17, 22, 23].
There are a few basic concepts that one must grasp in order to understand the
functionality and structure of Elasticsearch better.
Cluster Clusters are a collection of servers (nodes) that, when networked together,
hold the complete data set and provide federated indexing and search capability for
all connected servers.
Node A node is a solitary server that contains a portion of data and can contribute
to the cluster’s querying and indexing tasks. Note that each cluster is identified with
a name and a node can be assigned to a particular cluster based on the cluster’s
name. Starting an instance of Elasticsearch from scratch results in the creation of a
cluster with only one node. When the second instance of Elasticsearch starts (with
the same “cluster name”), the cluster will contain two nodes. Additional instances
Table 4.5 Relationship of
RDBMS terminology with
Elasticsearch
RDBMS Elasticsearch
Database Index
Table
Mapping
Field
Field
Row
JSON object
