Implementation and Performance Evaluation of Omni Compiler
75
Firstly, Omni compiler deletes a declaration of a local array a[][] and the
align directive. Next, Omni compiler creates a descriptor _XMP_DESC_a by a
function _XMP_init_array_desc() to set information of the distributed array.
Omni compiler also adds a function _XMP_alloc_array() to allocate memory
for the distributed array, and it sets values in an address _XMP_ADDR_a and a
leading dimension _XMP_ACC_a_0. Note that a multidimensional distributed array
is expressed as a one-dimensional array in the translated code since the size of each
dimension of the array may be determined dynamically.
2.2.2 Loop Statement
Figure 3 shows an XMP example code using a loop directive to parallelize the
following nested loop statement depending on the template t. Each dimension of
t is distributed onto two nodes, which is omitted there.
In the translated code above, a pointer _XMP_MULTI_ADDR_a is used which
has the size of each dimension as a head pointer of the distributed array a[][].
To improve performance, operations in a loop statement are performed using the
pointer [4]. Note that this pointer can be used when the number of elements in
each dimension of a distributed array is divisible by the number of nodes. If the
condition is not met, a one dimensional pointer _XMP_ADDR_a and an offset
_XMP_ACC_a_0 are used as shown in the translated code below.
Moreover, because values in ending conditions of the loop statement (i < 10, j <
10) are constants in a pre-translated code and are divisible by the number of nodes,
#pragma xmp loop on t[i][j]
for(int i=0;i<10;i++)
for(int j=0;j<10;j++)
a[i][j] = ...
double (*_XMP_MULTI_ADDR_a)[5] = (double (*)[5])(_XMP_ADDR_a);
for(int i=0;i<5;i++)
for(int j=0;j<5;j++)
_XMP_MULTI_ADDR_a[i][j] = ...
for(int i=0;i<5;i++)
for(int j=0;j<5;j++)
*(_XMP_ADDR_a + i * _XMP_ACC_a_0 + j) = ...
or
Fig. 3 Code translation of loop directive
75
Firstly, Omni compiler deletes a declaration of a local array a[][] and the
align directive. Next, Omni compiler creates a descriptor _XMP_DESC_a by a
function _XMP_init_array_desc() to set information of the distributed array.
Omni compiler also adds a function _XMP_alloc_array() to allocate memory
for the distributed array, and it sets values in an address _XMP_ADDR_a and a
leading dimension _XMP_ACC_a_0. Note that a multidimensional distributed array
is expressed as a one-dimensional array in the translated code since the size of each
dimension of the array may be determined dynamically.
2.2.2 Loop Statement
Figure 3 shows an XMP example code using a loop directive to parallelize the
following nested loop statement depending on the template t. Each dimension of
t is distributed onto two nodes, which is omitted there.
In the translated code above, a pointer _XMP_MULTI_ADDR_a is used which
has the size of each dimension as a head pointer of the distributed array a[][].
To improve performance, operations in a loop statement are performed using the
pointer [4]. Note that this pointer can be used when the number of elements in
each dimension of a distributed array is divisible by the number of nodes. If the
condition is not met, a one dimensional pointer _XMP_ADDR_a and an offset
_XMP_ACC_a_0 are used as shown in the translated code below.
Moreover, because values in ending conditions of the loop statement (i < 10, j <
10) are constants in a pre-translated code and are divisible by the number of nodes,
#pragma xmp loop on t[i][j]
for(int i=0;i<10;i++)
for(int j=0;j<10;j++)
a[i][j] = ...
double (*_XMP_MULTI_ADDR_a)[5] = (double (*)[5])(_XMP_ADDR_a);
for(int i=0;i<5;i++)
for(int j=0;j<5;j++)
_XMP_MULTI_ADDR_a[i][j] = ...
for(int i=0;i<5;i++)
for(int j=0;j<5;j++)
*(_XMP_ADDR_a + i * _XMP_ACC_a_0 + j) = ...
or
Fig. 3 Code translation of loop directive
