An efficient algorithm for out-of-core matrix transposition

doi:10.1109/12.995452

Journal ArticleDOI

An efficient algorithm for out-of-core matrix transposition

Jinwoo Suh, +1 more

- 01 Apr 2002 -

IEEE Transactions on Computers

- Vol. 51, Iss: 4, pp 420-438

TLDR

This paper proposes an algorithm that considers the index computation time and the I/O time and reduces the overall execution time and results in an overall reduction in the execution time due to the elimination of the expensive index computation.

Abstract:

Efficient transposition of out-of-core matrices has been widely studied. These efforts have focused on reducing the number of I/O operations. However, in state-of-the-art architectures, the memory-memory data transfer time and the index computation time are also significant components of the overall time. In this paper, we propose an algorithm that considers the index computation time and the I/O time and reduces the overall execution time. Our algorithm reduces the total execution time by reducing the number of I/O operations and eliminating the index computation. In doing so, two techniques are employed: writing the data on to disk in pre-defined patterns and balancing the number of disk read and write operations. The index computation time, which is an expensive operation involving two divisions and a multiplication, is eliminated by partitioning the memory into read and write buffers. The expensive in-processor permutation is replaced by data collection from the read buffer to the write buffer. Even though this partitioning may increase the number of I/O operations for some cases, it results in an overall reduction in the execution time due to the elimination of the expensive index computation. Our algorithm is analyzed using the well-known linear model and the parallel disk model. The experimental results on a Sun Enterprise, an SGI R12000 and a Pentium III show that our algorithm reduces the overall execution time by up to 50% compared with the best known algorithms in the literature.

An efficient algorithm for out-of-core matrix transposition

Citations

vLOD: high-fidelity walkthrough of large virtual environments

Generating SIMD vectorized permutations

Multidimensional signal processing

Efficient parallel out-of-core matrix transposition

Enhancing the matrix transpose operation using intel avx instruction set extension

References

Introduction to parallel computing: design and analysis of algorithms

RAID: high-performance, reliable secondary storage

The input/output complexity of sorting and related problems

Computability of Recursive Functions

Algorithms for parallel memory, I: Two-level memories

Related Papers (5)

Efficient transposition algorithms for large matrices

A Fast Computer Method for Matrix Transposing

Efficient parallel out-of-core matrix transposition

Parallel matrix transpose algorithms on distributed memory concurrent computers

VIS speeds new media processing