Scene labeling with LSTM recurrent neural networks

doi:10.1109/CVPR.2015.7298977

Proceedings ArticleDOI

Scene labeling with LSTM recurrent neural networks

Wonmin Byeon, +3 more

- pp 3547-3555

Chats0

TLDR

The approach, which has a much lower computational complexity than prior methods, achieved state-of-the-art performance over the Stanford Background and the SIFT Flow datasets and the ability to visualize feature maps from each layer supports the hypothesis that LSTM networks are overall suited for image processing tasks.

Abstract:

This paper addresses the problem of pixel-level segmentation and classification of scene images with an entirely learning-based approach using Long Short Term Memory (LSTM) recurrent neural networks, which are commonly used for sequence classification. We investigate two-dimensional (2D) LSTM networks for natural scene images taking into account the complex spatial dependencies of labels. Prior methods generally have required separate classification and image segmentation stages and/or pre- and post-processing. In our approach, classification, segmentation, and context integration are all carried out by 2D LSTM networks, allowing texture and spatial model parameters to be learned within a single model. The networks efficiently capture local and global contextual information over raw RGB values and adapt well for complex scene images. Our approach, which has a much lower computational complexity than prior methods, achieved state-of-the-art performance over the Stanford Background and the SIFT Flow datasets. In fact, if no pre- or post-processing is applied, LSTM networks outperform other state-of-the-art approaches. Hence, only with a single-core Central Processing Unit (CPU), the running time of our approach is equivalent or better than the compared state-of-the-art approaches which use a Graphics Processing Unit (GPU). Finally, our networks' ability to visualize feature maps from each layer supports the hypothesis that LSTM networks are overall suited for image processing tasks.

Scene labeling with LSTM recurrent neural networks

Citations

The Cityscapes Dataset for Semantic Urban Scene Understanding

Rethinking Atrous Convolution for Semantic Image Segmentation

Dual Attention Network for Scene Segmentation

The Cityscapes Dataset for Semantic Urban Scene Understanding

Object Detection With Deep Learning: A Review

References

ImageNet Classification with Deep Convolutional Neural Networks

Long short-term memory

Visualizing and Understanding Convolutional Networks

Backpropagation applied to handwritten zip code recognition

DeepFace: Closing the Gap to Human-Level Performance in Face Verification

Related Papers (5)

Long short-term memory

Fully convolutional networks for semantic segmentation

Deep Residual Learning for Image Recognition

Very Deep Convolutional Networks for Large-Scale Image Recognition

ImageNet Classification with Deep Convolutional Neural Networks