A
T WA
Â
à À1 A
T W ΔY ¼ ΔX
ð7Þ
W is a diagonal matrix containing a weight for each observation. The quantities in
the matrices on the left side are all known values, and the required shifts may be
computed using straightforward matrix multiplications and inversion.
1.4 Computing Power
For the three or four decades up to 2005, computer clock speeds increased exponentially each year from sub-Mhz until they reached approximately 2 GHz. This
trend supported by equivalent increase in power and capacity of other components
meant that code would just run faster each time it was ported to or installed on new
hardware. However, since 2005, clock speeds have stalled, and improvements have
been delivered by smaller transistors (Moore’s law) and multi-core processors. To
take advantage of these developments, software code often requires significant
reorganization to allow parts of substantial calculations to run in parallel.
High-performance mathematical libraries, in particular those based on BLAS and
LAPACK, make use of standardized cross-platform interfaces for linear algebra
operations which allows developers to focus on mathematical problems rather than
the computer science [14–16]. These libraries can outperform manual optimization
of code by several orders of magnitude and tend to scale to larger problems much
more efficiently. Most implementations also take advantage of multi-core processors
when available, speeding up code and saving developers from further work organizing the parallelization of calculations. Figure 3 illustrates an example of the
advantages of switching an existing least squares program to a high-performance
library: for a problem with 1,261 least squares parameters, during the computation of
Fig. 3 Comparison of matrix inversion time as a function of least squares problem size. Plot
generated data in reference [17]
Recent Developments in the Refinement and Analysis of Crystal Structures
51
T WA
Â
à À1 A
T W ΔY ¼ ΔX
ð7Þ
W is a diagonal matrix containing a weight for each observation. The quantities in
the matrices on the left side are all known values, and the required shifts may be
computed using straightforward matrix multiplications and inversion.
1.4 Computing Power
For the three or four decades up to 2005, computer clock speeds increased exponentially each year from sub-Mhz until they reached approximately 2 GHz. This
trend supported by equivalent increase in power and capacity of other components
meant that code would just run faster each time it was ported to or installed on new
hardware. However, since 2005, clock speeds have stalled, and improvements have
been delivered by smaller transistors (Moore’s law) and multi-core processors. To
take advantage of these developments, software code often requires significant
reorganization to allow parts of substantial calculations to run in parallel.
High-performance mathematical libraries, in particular those based on BLAS and
LAPACK, make use of standardized cross-platform interfaces for linear algebra
operations which allows developers to focus on mathematical problems rather than
the computer science [14–16]. These libraries can outperform manual optimization
of code by several orders of magnitude and tend to scale to larger problems much
more efficiently. Most implementations also take advantage of multi-core processors
when available, speeding up code and saving developers from further work organizing the parallelization of calculations. Figure 3 illustrates an example of the
advantages of switching an existing least squares program to a high-performance
library: for a problem with 1,261 least squares parameters, during the computation of
Fig. 3 Comparison of matrix inversion time as a function of least squares problem size. Plot
generated data in reference [17]
Recent Developments in the Refinement and Analysis of Crystal Structures
51
