18 Using Clustering of Panel Data to Examine Housing Demand …
187
also supported the study with SWOT analysis. Çelik and Kıral (2018b) used clustering of panel data and SWOT analysis to examined the socio-economic factors
affecting housing sales of the provinces of Turkey in the 2008–2015 process. They
determined factors affecting housing sales in provinces that exhibit similar characteristics in housing demand. The results obtained from the study showed that urbanization rate, the ratio of deposit interest rate, average household income, number
of household’s automobiles, stock market Istanbul 100 index, housing loan interest
rateangross return rate of housing, were significant in for housing demand. Akay
and Yüksel (2018) presented that the mixed panel dataset is clustered by agglomerative hierarchical algorithms based on Gower’s distance and by k-prototypes . Akay
and Yüksel (2019) suggested a new distance for clustering of the mixed variable
panel data set containing invariant time binary variable, without performing variable
conversion to avoid information loss.
18.3 Cluster Analysis of Panel Data with k-Prototype
Algorithm
Panel data refer to two-dimensional data which are obtained in time series and cross
section at the same time, and that means taking multiple cross sections on time series,
and selecting the sample observations on cross sections at the same time (Hou and Ai
2015). The poobility of the different topics in the data is one of the important issues
in the panel data. If the parameters in the regression can be considered homogeneous
between different subjects, different subjects can be brought together. However,
the normal situation is that subjects cannot be pooled due to high heterogeneity.
Some recent studies investigated the “partial poolability” by clustering subjects into
different groups so that subjects in the same cluster have homogeneous parameters
(Lu and Huang 2011).
Bonzo and Hermosilla (2002) applied probability link function to advance the
algorithm of the cluster, thus the cluster analysis could be effectively applied to the
analysis panel data. In this study, we choose k prototype algorithm to explain the
cluster analysis process of multivariable panel data.This algorithm was proposed
by Huang (1998). It is straightforward to integrate the k-modes and k-means algorithms into the k-prototypes algorithm used to cluster the mixed-type objects. Since
frequently encountered objects in real world databases are mixed-type objects, the kprototypes algorithm is practically more useful. The cost function is used in conjunction with a partitioned clustering algorithm. The cost function handles mixed datasets
and computes the distance between a data point and a centre of cluster in terms of
two distance values—one for the numeric attributes and the other for the categorical
attributes.The objective of k-prototype is to group the dataset X into k clusters by
minimizing the cost function,
Précédent

- 185/206

Suivant