6 Big Data
349
data processing is a very efficient way to process large amounts of data collected
over a period of time when there is no need for real-time analytics. Historically, this
has been the most common data processing approach. Traditional databases and data
warehouses, including Hadoop, are common examples of batch systems processing.
Stream processing typically utilizes continuous data and is a key component
in enabling fast data processing. Streaming enables almost instantaneously data
analysis of the data streaming from one device to another. This method of
continuous computation happens as data flows through the system with no required
time limitations on the output. Due to the near instant data flow, systems do not
require large amounts of data to be stored.
The streaming approach processes each new individual piece of data upon arrival.
In contrast to batch processing, there is no waiting until the next batch processing
interval. The term micro-batch is frequently associated with streaming, when
batches are small or processed at small intervals. Although processing may occur at
high frequency, data is still processed a batch at a time in the micro-batch paradigm.
Spark Streaming is an example of a system that supports micro-batch processing.
Stream processing is highly beneficial if the events are frequent, especially over
rapid time intervals, and there is a need for fast detection and response.
6.4.16 Spark Streaming
Spark Streaming is a Spark component that enables processing of live streams
of data by providing an API for manipulating data streams similar to Spark
Core’s RDD API. It enables scalable, high-throughput, fault-tolerant data stream
processing. Spark Streaming’s API enables the same high degree fault tolerance,
throughput, and scalability as Spark Core. Spark Streaming receives input data
streams and divides them into batches called DStreams. DStreams can be created
from a number of sources such as Kafka, Flume, and Kinesis or by applying
operations on other DStreams (Fig. 6.18).
6.4.17 Spark Functionality
Spark Streaming receives input data streams and divides the data into batches. These
batches are then processed by the Spark engine to generate the final stream of results
in batches.
Discretized Stream or DStream is the core concept enabled by Spark Streaming.
It represents a continuous stream of data. DStream is represented by a continuous
series of RDDs. Operations applied to DStreams translate to operations on the
underlying RDDs. Spark Streaming discretizes the data into small micro-batches.
Spark Streaming receivers accept data in parallel and buffer it in the workers nodes’
Précédent

- 354/647

Suivant