An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Open AccessProceedings Article

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

TLDR

The Vision Transformer (ViT) as discussed by the authors uses a pure transformer applied directly to sequences of image patches to perform very well on image classification tasks, achieving state-of-the-art results on ImageNet, CIFAR-100, VTAB, etc.

Abstract:

While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional networks, or used to replace certain components of convolutional networks while keeping their overall structure in place. We show that this reliance on CNNs is not necessary and a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks. When pre-trained on large amounts of data and transferred to multiple mid-sized or small image recognition benchmarks (ImageNet, CIFAR-100, VTAB, etc.), Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.

Citations

PDF

Open Access

More filters

Posted Content

An Empirical Study of Training Self-Supervised Vision Transformers

Xinlei Chen, +2 more

- 05 Apr 2021 -

arXiv: Computer Vision and Pattern Recog...

TL;DR: This work investigates the effects of several fundamental components for training self-supervised ViT, and reveals that these results are indeed partial failure, and they can be improved when training is made more stable.

...read moreread less

Posted Content

Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Wenhai Wang, +8 more

- 24 Feb 2021 -

arXiv: Computer Vision and Pattern Recog...

TL;DR: Huang et al. as discussed by the authors proposed Pyramid Vision Transformer (PVT), which is a simple backbone network useful for many dense prediction tasks without convolutions, and achieved state-of-the-art performance on the COCO dataset.

...read moreread less

Posted Content

Natural Adversarial Examples

Dan Hendrycks, +4 more

- 16 Jul 2019 -

arXiv: Learning

TL;DR: This work introduces two challenging datasets that reliably cause machine learning model performance to substantially degrade and curates an adversarial out-of-distribution detection dataset called IMAGENET-O, which is the first out- of-dist distribution detection dataset created for ImageNet models.

...read moreread less

Posted Content

An Attentive Survey of Attention Models

Sneha Chaudhari, +3 more

- 05 Apr 2019 -

arXiv: Learning

TL;DR: A taxonomy that groups existing techniques into coherent categories in attention models is proposed, and how attention has been used to improve the interpretability of neural networks is described.

...read moreread less

Posted Content

CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification

Chun-Fu Chen, +2 more

- 27 Mar 2021 -

arXiv: Computer Vision and Pattern Recog...

TL;DR: Zhang et al. as mentioned in this paper proposed a dual-branch transformer to combine image patches (i.e., tokens in a transformer) of different sizes to produce stronger image features, which achieved promising results on image classification compared to convolutional neural networks.

...read moreread less

Collapse

References

PDF

Open Access

More filters

Posted Content

Language Models are Few-Shot Learners

Tom B. Brown, +30 more

- 28 May 2020 -

arXiv: Computation and Language

TL;DR: This article showed that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches.

...read moreread less

Proceedings ArticleDOI

Self-Training With Noisy Student Improves ImageNet Classification

Qizhe Xie, +3 more

TL;DR: A simple self-training method that achieves 88.4% top-1 accuracy on ImageNet, which is 2.0% better than the state-of-the-art model that requires 3.5B weakly labeled Instagram images.

...read moreread less

Proceedings Article

A Simple Framework for Contrastive Learning of Visual Representations

Ting Chen, +3 more

TL;DR: SimCLR as mentioned in this paper is a simple framework for contrastive learning of visual representations and achieves state-of-the-art performance on ImageNet. But it requires large batch sizes and more training steps compared to supervised learning.

...read moreread less

Book ChapterDOI

UNITER: UNiversal Image-TExt Representation Learning

Yen-Chun Chen, +7 more

TL;DR: UNITER, a UNiversal Image-TExt Representation, learned through large-scale pre-training over four image-text datasets is introduced, which can power heterogeneous downstream V+L tasks with joint multimodal embeddings.

...read moreread less

Proceedings ArticleDOI

CCNet: Criss-Cross Attention for Semantic Segmentation

Zilong Huang, +5 more

TL;DR: CCNet as mentioned in this paper proposes a recurrent criss-cross attention module to harvest the contextual information of all the pixels on its crisscross path, and then takes a further recurrent operation to finally capture the full-image dependencies from all pixels.

...read moreread less

Collapse

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Citations

An Empirical Study of Training Self-Supervised Vision Transformers

Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Natural Adversarial Examples

An Attentive Survey of Attention Models

CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification

References

Language Models are Few-Shot Learners

Self-Training With Noisy Student Improves ImageNet Classification

A Simple Framework for Contrastive Learning of Visual Representations

UNITER: UNiversal Image-TExt Representation Learning

CCNet: Criss-Cross Attention for Semantic Segmentation

Related Papers (5)

Attention is All you Need

Deep Residual Learning for Image Recognition

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

ImageNet: A large-scale hierarchical image database

Microsoft COCO: Common Objects in Context