316
N. Balac
6.4.4 RDD Transformations and Actions ............................................. 338
6.4.5 Datasets and DataFrames in Spark .............................................. 340
6.4.6 The Spark Processing Engine ................................................... 340
6.4.7 Spark Components ............................................................... 341
6.4.8 Spark SQL........................................................................ 343
6.4.9 Spark DataFrames ............................................................... 344
6.4.10 Creating a DataFrame ........................................................... 344
6.4.11 DataFrame Operations ........................................................... 345
6.4.12 Spark MLlib ...................................................................... 346
6.4.13 MLlib Capabilities ............................................................... 347
6.4.14 Spark Streaming ................................................................. 348
6.4.15 Intro to Batch and Stream Processing........................................... 348
6.4.16 Spark Streaming ................................................................. 349
6.4.17 Spark Functionality .............................................................. 349
6.5 Big Data Analytics: Building the Data Pipeline ......................................... 352
6.5.1 Developing Predictive and Prescriptive Models ................................ 352
6.5.2 The Cross Industry Standard Process for Data Mining (CRISP-DM) ......... 353
6.6 Conclusion................................................................................. 354
References ........................................................................................ 354
6.1 Introduction to Big Data
Big Data has changed the way we manage, analyze, and leverage data across
all industry sectors. Big Data has the potential to examine and reveal trends,
find unseen patterns, discover hidden correlations, reveal new information, extract
insight, enhance decision making and automation, etc. Managing and analyzing
data, in particular in the era of IoT, has always been one of the greatest challenges
within organizations across industries. Finding an efficient and scalable approach
to capturing, integrating, organizing, and analyzing information about IoT devices,
products, and services can be a perplexing task for any organization regardless of
the size or line of business. In the age of the Internet and digital transformation,
the notion of Big Data reflects the changing world we live in. Everywhere around
the world, more and more data is captured and recorded every day. Companies and
organizations are becoming overwhelmed by the complexity and sheer volume of
their data. While some data is still structured and stored in a traditional relational
databases or data warehouses, a vast majority of the modern data sources are
producing unstructured data including documents, conversations, pictures, videos,
Tweets, posts, Snapchats, sensor readouts, click streams, and machine-to-machine
data. Further, the availability and adoption of newer, more powerful mobile devices,
coupled with ubiquitous access to global networks, is continuously driving the
creation of new sources for data.
Although each data source can be independently managed and searched, the
challenge today is how to make sense of the intersection of all these different
types of data. When large amounts of data are coming from so many different
forms, traditional data management techniques falter. While there has always been
a challenging amount of data for existing IT infrastructure, the difference today is
Précédent

- 321/647

Suivant