Eigen Analysis of Self-Attention and its Reconstruction from Partial Computation.

Open AccessPosted Content

Eigen Analysis of Self-Attention and its Reconstruction from Partial Computation.

Srinadh Bhojanapalli, +5 more

- 16 Jun 2021 -

arXiv: Learning

Chats0

TLDR

In this article, the authors investigate the global structure of attention scores computed using this dot product mechanism on a typical distribution of inputs, and study the principal components of their variation through eigen analysis of full attention score matrices.

Abstract:

State-of-the-art transformer models use pairwise dot-product based self-attention, which comes at a computational cost quadratic in the input sequence length. In this paper, we investigate the global structure of attention scores computed using this dot product mechanism on a typical distribution of inputs, and study the principal components of their variation. Through eigen analysis of full attention score matrices, as well as of their individual rows, we find that most of the variation among attention scores lie in a low-dimensional eigenspace. Moreover, we find significant overlap between these eigenspaces for different layers and even different transformer models. Based on this, we propose to compute scores only for a partial subset of token pairs, and use them to estimate scores for the remaining pairs. Beyond investigating the accuracy of reconstructing attention scores themselves, we investigate training transformer models that employ these approximations, and analyze the effect on overall accuracy. Our analysis and the proposed method provide insights into how to balance the benefits of exact pair-wise attention and its significant computational expense.

Eigen Analysis of Self-Attention and its Reconstruction from Partial Computation.

Citations

Leveraging redundancy in attention with Reuse Transformers.

References

Attention is All you Need

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference

GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Related Papers (5)

Conditional out-of-sample generation for unpaired data using trVAE.

Selecting the number of components in PCA via random signflips

Information theoretic similarity measures for shape matching

Adjusting Leverage Scores by Row Weighting: A Practical Approach to Coherent Matrix Completion

Sparse Functional Principal Component Analysis in High Dimensions