run without flushing the output buffers prompt]$ mpirun -np 8 prefix_sum At step 1, node 0 has number 1. At step 2, node 0 has number 1. At step 3, node 0 has number 1. At step 1, Node 2 has number 5 = 3 + 2. At step 2, Node 2 has number 6 = 5 + 1. At step 3, node 2 has number 6. At step 1, Node 6 has number 13 = 7 + 6. At step 2, Node 6 has number 22 = 13 + 9. At step 3, Node 6 has number 28 = 22 + 6. At step 1, Node 1 has number 3 = 2 + 1. At step 2, node 1 has number 3. At step 3, node 1 has number 3. At step 1, Node 3 has number 7 = 4 + 3. At step 2, Node 3 has number 10 = 7 + 3. At step 3, node 3 has number 10. At step 1, Node 4 has number 9 = 5 + 4. At step 2, Node 4 has number 14 = 9 + 5. At step 3, Node 4 has number 15 = 14 + 1. At step 1, Node 5 has number 11 = 6 + 5. At step 2, Node 5 has number 18 = 11 + 7. At step 3, Node 5 has number 21 = 18 + 3. At step 1, Node 7 has number 15 = 8 + 7. At step 2, Node 7 has number 26 = 15 + 11. At step 3, Node 7 has number 36 = 26 + 10. The total sum is 36. prompt]$ run with flushing the output buffers prompt]$ mpirun -np 8 prefix_sum At step 1, node 0 has number 1. At step 1, Node 4 has number 9 = 5 + 4. At step 1, Node 3 has number 7 = 4 + 3. At step 1, Node 2 has number 5 = 3 + 2. At step 1, Node 6 has number 13 = 7 + 6. At step 1, Node 7 has number 15 = 8 + 7. At step 1, Node 1 has number 3 = 2 + 1. At step 1, Node 5 has number 11 = 6 + 5. At step 2, node 1 has number 3. At step 2, Node 5 has number 18 = 11 + 7. At step 2, Node 7 has number 26 = 15 + 11. At step 2, Node 3 has number 10 = 7 + 3. At step 2, Node 2 has number 6 = 5 + 1. At step 2, Node 4 has number 14 = 9 + 5. At step 2, Node 6 has number 22 = 13 + 9. At step 2, node 0 has number 1. At step 3, node 0 has number 1. At step 3, Node 4 has number 15 = 14 + 1. At step 3, Node 5 has number 21 = 18 + 3. At step 3, node 1 has number 3. At step 3, node 3 has number 10. At step 3, node 2 has number 6. At step 3, Node 6 has number 28 = 22 + 6. At step 3, Node 7 has number 36 = 26 + 10. The total sum is 36. prompt]$For p processors, the algorithm finishes in log2(p) steps, and thus achieves a respectable computational speedup, although every "+" needs one send and one recv operation, so the communication overhead will dominate. As an application for this prefix sum algorithm we recalled the composite trapezoidal rule, here for this case we take 8 function evaluations using 7 subintervals. As the computation time is now dominated by the function evaluations, we may expect the same good speedup as of any embarrassingly parallel computation. We concluded by stating that the prefix sum algorithm can be seen as an efficient implementation of the MPI_Reduce command which sums up the results from all nodes.
As an example of synchronous iteration, we outlined a parallel version of the Jacobi method to solve a linear system iteratively. At every iteration we need to invoke a barrier. We ended mentioning the application of the butterfly barrier to implement the stop criterium.
Bibliography