XcalableACC: An Integration of XcalableMP and OpenACC
129
Fig. 5 XACC code with OpenACC loop construct
to accelerator memory. In lines 8–9, the parallel directive and XMP loop
directive parallelize the following loop on an accelerator within a node and among
nodes, respectively.
Example 2
In lines 2–5 of Fig. 6, the directives declare a global array a. In line 7, the
copy clause on the parallel directive transfers a and a variable sum from
host memory to accelerator memory. In lines 7–8, the parallel directive and
XMP loop directive parallelize the following loop on an accelerator within a node
and in among nodes, respectively. After finishing the calculation of the loop, the
OpenACC reduction clause and the XMP reduction clause with acc in lines
7–8 perform a reduction operation for sum first on the accelerator within a node and
then among all nodes.
2.3 Data Communication and Synchronization
When an acc clause is specified in an XMP’s communication and synchronization
directive, the directive works for the data on accelerator memory to transfer it.
The acc clause can be specified on the following XMP’s communication and
synchronization directives:
• reflect
• gmove
• barrier
• reduction
• bcast
• wait_async
Précédent

- 136/265

Suivant