218
F. Firouzi and B. Farahani
Table 4.7 Common file formats in data lake
File format Properties
Use cases
JSON
Flexible and human
readable
The consumer of data is an application which
needs small data volume
CSV
Fixed schema, but human
readable
The consumer is an analysis and the data volume
is small
Orc
Columnar with schema
Efficient compression and improved performance
for reading, writing, and processing data, e.g.,
suitable for read-heavy analytical loads
Parquet
Columnar with schema
which supports complex
nested data structures
Designed for projects in the Hadoop ecosystem.
Suitable for read-heavy analytical loads
Avro
Row-based with schema
Suitable for write-heavy workloads
Service-oriented architecture
(coarse-grained)
Microservices
(fine-grained)
Monolithic
(one unit)
Fig. 4.31 Engineering server-side IoT applications via three different software architectures
zones based on the specific use case. For example, a tier could be separated
into a warehouse staging zone, a machine learning training data zone, etc. Every
tier/zone may include its own file formats and individual, granular access levels.
The well-known file formats, their properties, and use cases are demonstrated in
Table 4.7. Finally, note that transferring data between data lake tiers/zones involves
a computing resource to complete the data transformation and process data between
tiers/zones [27].
4.7 Application Layer
IoT applications located in the Cloud can be engineered based on three distinct
architectures, namely, (i) monolithic architecture, (ii) service-oriented architecture
(SOA), and (iii) microservice architecture. Figure 4.31 demonstrates the general
differences between these three options [28].
When creating a server-side application, you can begin with a layered architecture. This architecture typically consists of the following layers/tiers [28]:
F. Firouzi and B. Farahani
Table 4.7 Common file formats in data lake
File format Properties
Use cases
JSON
Flexible and human
readable
The consumer of data is an application which
needs small data volume
CSV
Fixed schema, but human
readable
The consumer is an analysis and the data volume
is small
Orc
Columnar with schema
Efficient compression and improved performance
for reading, writing, and processing data, e.g.,
suitable for read-heavy analytical loads
Parquet
Columnar with schema
which supports complex
nested data structures
Designed for projects in the Hadoop ecosystem.
Suitable for read-heavy analytical loads
Avro
Row-based with schema
Suitable for write-heavy workloads
Service-oriented architecture
(coarse-grained)
Microservices
(fine-grained)
Monolithic
(one unit)
Fig. 4.31 Engineering server-side IoT applications via three different software architectures
zones based on the specific use case. For example, a tier could be separated
into a warehouse staging zone, a machine learning training data zone, etc. Every
tier/zone may include its own file formats and individual, granular access levels.
The well-known file formats, their properties, and use cases are demonstrated in
Table 4.7. Finally, note that transferring data between data lake tiers/zones involves
a computing resource to complete the data transformation and process data between
tiers/zones [27].
4.7 Application Layer
IoT applications located in the Cloud can be engineered based on three distinct
architectures, namely, (i) monolithic architecture, (ii) service-oriented architecture
(SOA), and (iii) microservice architecture. Figure 4.31 demonstrates the general
differences between these three options [28].
When creating a server-side application, you can begin with a layered architecture. This architecture typically consists of the following layers/tiers [28]:
