Parallel Numerical Linear Algebra
We like to compute at processor speed, not at memory speed.
Our goal is to restrict the data movement to O(n^2)
while doing as many floating point operations as O(n^3).
In this lecture, we sketched the 3-level structure of the BLAS
(Basic Linear Algebra Subprograms) library of
LAPACK.
We then described the organization of LU factorization in blocked form,
distinguishing between the "right looking LU" (which is similar to
how we normally do LU) and the "left looking LU".
The latter is better for data movement.
Bibliography
- James W. Demmel: "Applied Numerical Linear Algebra", SIAM 1997.
- Jack Dongarra, Iain S. Duff, Danny C. Sorensen,
and Henk A. van der Vorst:
"Solving Linear Systems on Vector and Shared Memory Computers",
SIAM 1991.