56
R. Lulli et al.
Fig. 7.5 Comparison chart of the performance in FFT computation time of different Nvidia GPU
boards with increasing data set dimensions (one test every million of points)
have) an high cost that exceeded our limited budget, so we were forced to substitute
Tesla boards with high-end products of the entertainment dedicated GeForce and
Titan line.
A summary chart of the most significant Nvidia GPUs tests we carried out is in
Fig. 7.5. In the tests, we used four available high-end GPUs different as in chipset
technology as in available onboard RAM:
• Quadro 5000, an industry-standard 40 nm Fermi series GPU of 2010 with 352
cores and 2.5 GB Graphics Double Data Rate 5 (GDDR5) [28];
• GeForce GTX 680, an entertainment 28 nm Kepler series GPU of 2012 with 1536
cores and 4 GB GDDR5 [29];
• GeForce GTX 970, an entertainment 28 nm Maxwell series GPU of 2014 with
1664 cores and 4 GB GDDR5 [30];
• TitanXp a top-end entertainment 16 nm Pascal series GPU of 2017 with 3840
cores and 12 GB GDDR5 [31].
With each GPU, we computed the FFT of a big single block data set produced by
the Ultraview AD8-1500 × 2 DAQ board, so in the tests, we were limited only by
GPU available onboard memory. With the algorithm we used, we realized that the
computation was always completed in less than 350 ms having spare time both for
other computations and for data movement to and from the accelerator.
R. Lulli et al.
Fig. 7.5 Comparison chart of the performance in FFT computation time of different Nvidia GPU
boards with increasing data set dimensions (one test every million of points)
have) an high cost that exceeded our limited budget, so we were forced to substitute
Tesla boards with high-end products of the entertainment dedicated GeForce and
Titan line.
A summary chart of the most significant Nvidia GPUs tests we carried out is in
Fig. 7.5. In the tests, we used four available high-end GPUs different as in chipset
technology as in available onboard RAM:
• Quadro 5000, an industry-standard 40 nm Fermi series GPU of 2010 with 352
cores and 2.5 GB Graphics Double Data Rate 5 (GDDR5) [28];
• GeForce GTX 680, an entertainment 28 nm Kepler series GPU of 2012 with 1536
cores and 4 GB GDDR5 [29];
• GeForce GTX 970, an entertainment 28 nm Maxwell series GPU of 2014 with
1664 cores and 4 GB GDDR5 [30];
• TitanXp a top-end entertainment 16 nm Pascal series GPU of 2017 with 3840
cores and 12 GB GDDR5 [31].
With each GPU, we computed the FFT of a big single block data set produced by
the Ultraview AD8-1500 × 2 DAQ board, so in the tests, we were limited only by
GPU available onboard memory. With the algorithm we used, we realized that the
computation was always completed in less than 350 ms having spare time both for
other computations and for data movement to and from the accelerator.
