(in the form of a list) specifying the types of some or all variables in the dataset.
When there are missing values (coded NA) in the data, the function excludes from
the comparison of two sites a variable where one or the other object (or both) has a
missing value.
gowdis() of package FD is the most complete function to compute Gower’s
coefficient. It computes the distance for mixed variables including asymmetrical
binary variables. It provides three ways of handling ordinal variables, including
the method of Podani (1999). It offers the same treatment of missing values as in
daisy() and allows users to attribute different weights to the variables.
Let us again create an artificial dataset containing four variables: two random
quantitative variables and two factors:
# Fictitious data for Gower (S15) index
# Random normal deviates with zero mean and unit standard deviation
var.g1 <- rnorm(30, 0, 1)
# Random uniform deviates from 0 to 5
var.g2 <- runif(30, 0, 5)
# Factor with 3 levels (10 objects each)
var.g3 <- gl(3, 10, labels = c("A", "B", "C"))
# Factor with 2 levels, orthogonal to var.g3
var.g4 <- gl(2, 5, 30, labels = c("D", "E"))
Together, var.g3 and var.g4 represent a 2-way crossed balanced design.
dat2 <- data.frame(var.g1, var.g2, var.g3, var.g4)
summary(dat2)
Hints Function gl() is quite handy to generate factors, but by default it uses numerals
as labels for the levels. Use argument labels to provide alphanumeric
characters instead of numbers.
Note the use of data.frame() to assemble the four variables. Unlike
cbind(), data.frame() preserves the classes of the variables. Variables 3
and 4 thus retain their class "factor".
Let us first compute and view the complete S 15 matrix. Then repeat the computation using the two factors (var.g3 and var.g4) only:
50
3 Association Measures and Matrices
Précédent

- 64/444

Suivant