8.2.2 State-of-the-Art Analysis Methods
The purpose of big data analytics is to derive useful information and knowledge
from big data, while the main purpose of the big data management is to make big
data analytics possible and feasible. In traditional data analysis, causal relationship
among the variables is normally sought from data samples. Because of the number of
variables, the volume of data and the uncertainty in the data quality involved in the
big data analytics, correlation, or probability relationship are usually sought by
analyzing whole datasets instead of samples.
Recently, a number of popular big data processing platforms come in the market
and draw a lot of attention. Google Earth Engine (GEE) is the most frequently
mentioned name of web service platforms in the geoinformatics and agricultural
information communities (Gorelick et al. 2017). Living in the vast Google Cloud,
Google Earth Engine serves planetary-scale geospatial analysis capability, which
brings Google’s massive computational capability to solve high-impact scientific
problems like deforestation, drought, disaster, disease, food security, water management, climate change, and environment protection. It gives scientists a free yet very
powerful tool to deal with big data. Because GEE integrates geospatial big data,
analytics algorithms and models, and powerful computing capability into a single
platform, it enables many big data applications, saves tremendous amount of time for
scientists, and promotes research democracy. Currently, most remote sensing data,
but not ground truth data, used in agro-geoinformatics are available in GEE.
For agro-big data controlled by individual scientists or organizations, they have to
maintain their own big data storage and computing and analysis environments. There
are several open source software that scientists used frequently, such as Eucalptus,
Openstack (Sefraoui et al. 2012), Cloudstack (Kumar et al. 2014), Apache Hadoop,
and Apache Spark. These software can create a cloud-based environment and
accelerate the big data processing speed by the MapReduce mechanism (Borthakur
2007; Zaharia et al. 2016). Due to the five V challenges, these software ecosystems
often run into constant issues dealing with bottleneck problems like data transferring
among hosts and memory, MapReduce optimization, real-time data processing,
system failure tolerances, data security, etc. The applied analysis methods are also
changing. Scientists used to run traditional numeric models of the crop growth and
predict the future development of crop by simulations. In recent years, agricultural
scientists start to use artificial intelligence technology like machine learning for more
effortless, low-cost, general, and automatic modeling of the crops and surrounding
environment (Yu et al. 2018; Sun et al. 2019).
148
L. Di and Z. Sun
Précédent

- 152/419

Suivant