4 Architecting IoT Cloud
183
to real-time streaming, data is ingested as it arrives. On the other hand, in batch
mode, the data is ingested in pieces at periodic time intervals. As the availability of
IoT devices grows, the variety and volume of data sources are quickly expanding.
Therefore, obtaining data can be a challenge. In general, major challenges facing
data ingestion include [5, 6]:
• Velocity and Volume of Data – Data volume and the frequency of data generation
are incredibly high in IoT.
• Heterologous Data Sources – When data sources have different formats and
natures, it can be difficult to ingest quickly, process, and prioritize data.
• Rapid Evolution – Both data sources and ingestion technologies/frameworks
change quickly over time.
• Independent Data Changes – Data can change independently from the ingestion
application without any prior notice.
• Semantic Data Changes – Data semantics can evolve as the data is used in new
business scenarios.
• Detection and Capture – Detecting and obtaining altered data can be challenging
due to the unstructured or semi-structured nature of data as well as the lowlatency requirements of specific business cases.
When designing a Data Ingestion System, it is important to consider the
following:
• Upgrade Capability – The system must be able to upgrade in order to handle new
data sources, applications, and technologies.
• Data Integrity – We have to ensure that the ingestion application is consistently
obtaining correct and trustworthy data.
• Dependability and Fault Tolerance – The system must be fault tolerant of
overcoming any failures.
• Data Volume – The ingestion layer should be able to handle the high volume
of data. In general, the ability to store all data is preferable; however, in some
instances it may be more appropriate to store aggregated (processed) data.
• Scalability – The system must rapidly consume data, be able to scale based on
volume as well as the speed at which big data of IoT comes in from various
sources including networks, machinery, sensors, human interaction, social media,
and other media sites.
• Heterogeneous Data Source/Format – While data can take different forms, it is
usually structured (i.e., tabular one), unstructured (i.e., video, audio, images),
or semi-structured (i.e., CSS files, JSON files, etc.). The data ingestion layer
should be capable of utilizing various data sources, data formats, technologies,
and operating systems.
183
to real-time streaming, data is ingested as it arrives. On the other hand, in batch
mode, the data is ingested in pieces at periodic time intervals. As the availability of
IoT devices grows, the variety and volume of data sources are quickly expanding.
Therefore, obtaining data can be a challenge. In general, major challenges facing
data ingestion include [5, 6]:
• Velocity and Volume of Data – Data volume and the frequency of data generation
are incredibly high in IoT.
• Heterologous Data Sources – When data sources have different formats and
natures, it can be difficult to ingest quickly, process, and prioritize data.
• Rapid Evolution – Both data sources and ingestion technologies/frameworks
change quickly over time.
• Independent Data Changes – Data can change independently from the ingestion
application without any prior notice.
• Semantic Data Changes – Data semantics can evolve as the data is used in new
business scenarios.
• Detection and Capture – Detecting and obtaining altered data can be challenging
due to the unstructured or semi-structured nature of data as well as the lowlatency requirements of specific business cases.
When designing a Data Ingestion System, it is important to consider the
following:
• Upgrade Capability – The system must be able to upgrade in order to handle new
data sources, applications, and technologies.
• Data Integrity – We have to ensure that the ingestion application is consistently
obtaining correct and trustworthy data.
• Dependability and Fault Tolerance – The system must be fault tolerant of
overcoming any failures.
• Data Volume – The ingestion layer should be able to handle the high volume
of data. In general, the ability to store all data is preferable; however, in some
instances it may be more appropriate to store aggregated (processed) data.
• Scalability – The system must rapidly consume data, be able to scale based on
volume as well as the speed at which big data of IoT comes in from various
sources including networks, machinery, sensors, human interaction, social media,
and other media sites.
• Heterogeneous Data Source/Format – While data can take different forms, it is
usually structured (i.e., tabular one), unstructured (i.e., video, audio, images),
or semi-structured (i.e., CSS files, JSON files, etc.). The data ingestion layer
should be capable of utilizing various data sources, data formats, technologies,
and operating systems.
