Online Sequence Training of Recurrent Neural Networks with Connectionist Temporal Classification.

Open AccessPosted Content

Online Sequence Training of Recurrent Neural Networks with Connectionist Temporal Classification.

- 21 Nov 2015 -

TLDR

An expectation-maximization (EM) based online CTC algorithm is introduced that enables unidirectional RNNs to learn sequences that are longer than the amount of unrolling and can also be trained to process an infinitely long input sequence without pre-segmentation or external reset.

Abstract:

Connectionist temporal classification (CTC) based supervised sequence training of recurrent neural networks (RNNs) has shown great success in many machine learning areas including end-to-end speech and handwritten character recognition. For the CTC training, however, it is required to unroll (or unfold) the RNN by the length of an input sequence. This unrolling requires a lot of memory and hinders a small footprint implementation of online learning or adaptation. Furthermore, the length of training sequences is usually not uniform, which makes parallel training with multiple sequences inefficient on shared memory models such as graphics processing units (GPUs). In this work, we introduce an expectation-maximization (EM) based online CTC algorithm that enables unidirectional RNNs to learn sequences that are longer than the amount of unrolling. The RNNs can also be trained to process an infinitely long input sequence without pre-segmentation or external reset. Moreover, the proposed approach allows efficient parallel training on GPUs. For evaluation, phoneme recognition and end-to-end speech recognition examples are presented on the TIMIT and Wall Street Journal (WSJ) corpora, respectively. Our online model achieves 20.7% phoneme error rate (PER) on the very long input sequence that is generated by concatenating all 192 utterances in the TIMIT core test set. On WSJ, a network can be trained with only 64 times of unrolling while sacrificing 4.5% relative word error rate (WER).

Online Sequence Training of Recurrent Neural Networks with Connectionist Temporal Classification.

Citations

Deep Lip Reading: a comparison of models and an online application

Character-level incremental speech recognition with recurrent neural networks

Online Keyword Spotting with a Character-Level Recurrent Neural Network.

Influenza-like illness prediction using a long short-term memory deep learning model with multiple open data sources

An End-to-end Framework for Audio-to-Score Music Transcription on Monophonic Excerpts

References

Long short-term memory

Neural Machine Translation by Jointly Learning to Align and Translate

Learning Phrase Representations using RNN Encoder--Decoder for Statistical Machine Translation

Neural Machine Translation by Jointly Learning to Align and Translate

Sequence to Sequence Learning with Neural Networks

Related Papers (5)

Self-attention Networks for Connectionist Temporal Classification in Speech Recognition

Connectionist Temporal Classification

Sequence training of multiple deep neural networks for better performance and faster training speed

Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer

Handwritten Digit String Recognition by Combination of Residual Network and RNN-CTC