Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering

doi:10.1109/CVPR.2018.00636

Open AccessProceedings ArticleDOI

Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering

Peter Anderson, +6 more

- pp 6077-6086

Chats0

TLDR

In this paper, a bottom-up and top-down attention mechanism was proposed to enable attention to be calculated at the level of objects and other salient image regions, which achieved state-of-the-art results on the MSCOCO test server.

Abstract:

Top-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained analysis and even multiple steps of reasoning. In this work, we propose a combined bottom-up and top-down attention mechanism that enables attention to be calculated at the level of objects and other salient image regions. This is the natural basis for attention to be considered. Within our approach, the bottom-up mechanism (based on Faster R-CNN) proposes image regions, each with an associated feature vector, while the top-down mechanism determines feature weightings. Applying this approach to image captioning, our results on the MSCOCO test server establish a new state-of-the-art for the task, achieving CIDEr / SPICE / BLEU-4 scores of 117.9, 21.5 and 36.9, respectively. Demonstrating the broad applicability of the method, applying the same approach to VQA we obtain first place in the 2017 VQA Challenge.

Citations

PDF

Open Access

More filters

Posted Content

Answer Them All! Toward Universal Visual Question Answering Models

Robik Shrestha, +2 more

- 01 Mar 2019 -

arXiv: Computer Vision and Pattern Recog...

TL;DR: In this paper, the authors compare five state-of-the-art VQA algorithms across eight different datasets covering both natural image understanding and synthetic data sets and find that methods do not generalize across the two domains.

...read moreread less

Book ChapterDOI

Learning to Generate Grounded Visual Captions Without Localization Supervision

Chih-Yao Ma, +5 more

TL;DR: This paper proposed a cyclical training regimen that forces the model to localize each word in the image after the sentence decoder generates it, and then reconstruct the sentence from the localized image region(s) to match the ground-truth.

...read moreread less

Posted Content

Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCaps

Qi Zhu, +3 more

- 09 Dec 2020 -

arXiv: Computer Vision and Pattern Recog...

TL;DR: This paper argues that a simple attention mechanism can do the same or even better job without any bells and whistles of multi-modality encoder design, and finds this simple baseline model consistently outperforms state-of-the-art (SOTA) models on two popular benchmarks, TextVQA and all three tasks of ST-V QA.

...read moreread less

Proceedings ArticleDOI

Object-Difference Attention: A Simple Relational Attention for Visual Question Answering

Chenfei Wu, +3 more

TL;DR: An object-difference attention (ODA) is proposed which calculates the probability of attention by implementing difference operator between different image objects in an image under the guidance of questions in hand and Experimental results show those relational attentions have strengths on different types of questions.

...read moreread less

Proceedings ArticleDOI

Cascade Reasoning Network for Text-based Visual Question Answering

Fen Liu, +5 more

TL;DR: A novel Cascade Reasoning Network (CRN) is proposed that consists of a progressive attention module (PAM) and a multimodal reasoning graph (MRG) module that aims to explicitly model the connections and interactions between texts and visual concepts.

...read moreread less

Collapse

References

PDF

Open Access

More filters

Proceedings ArticleDOI

Deep Residual Learning for Image Recognition

Kaiming He, +3 more

TL;DR: In this article, the authors proposed a residual learning framework to ease the training of networks that are substantially deeper than those used previously, which won the 1st place on the ILSVRC 2015 classification task.

...read moreread less

Journal ArticleDOI

Long short-term memory

Sepp Hochreiter, +1 more

- 01 Nov 1997 -

Neural Computation

TL;DR: A novel, efficient, gradient based method called long short-term memory (LSTM) is introduced, which can learn to bridge minimal time lags in excess of 1000 discrete-time steps by enforcing constant error flow through constant error carousels within special units.

...read moreread less

Journal ArticleDOI

ImageNet Large Scale Visual Recognition Challenge

Olga Russakovsky, +11 more

- 01 Dec 2015 -

International Journal of Computer Vision

TL;DR: The ImageNet Large Scale Visual Recognition Challenge (ILSVRC) as mentioned in this paper is a benchmark in object category classification and detection on hundreds of object categories and millions of images, which has been run annually from 2010 to present, attracting participation from more than fifty institutions.

...read moreread less

Book ChapterDOI

Microsoft COCO: Common Objects in Context

Tsung-Yi Lin, +7 more

TL;DR: A new dataset with the goal of advancing the state-of-the-art in object recognition by placing the question of object recognition in the context of the broader question of scene understanding by gathering images of complex everyday scenes containing common objects in their natural context.

...read moreread less

Proceedings ArticleDOI

You Only Look Once: Unified, Real-Time Object Detection

Joseph Redmon, +3 more

TL;DR: Compared to state-of-the-art detection systems, YOLO makes more localization errors but is less likely to predict false positives on background, and outperforms other detection methods, including DPM and R-CNN, when generalizing from natural images to other domains like artwork.

...read moreread less

Collapse

Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering

Citations

Answer Them All! Toward Universal Visual Question Answering Models

Learning to Generate Grounded Visual Captions Without Localization Supervision

Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCaps

Object-Difference Attention: A Simple Relational Attention for Visual Question Answering

Cascade Reasoning Network for Text-based Visual Question Answering

References

Deep Residual Learning for Image Recognition

Long short-term memory

ImageNet Large Scale Visual Recognition Challenge

Microsoft COCO: Common Objects in Context

You Only Look Once: Unified, Real-Time Object Detection

Related Papers (5)

Deep Residual Learning for Image Recognition

Microsoft COCO: Common Objects in Context

Attention is All you Need

Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

Faster R-CNN: towards real-time object detection with region proposal networks