Semantic Sentence Matching with Densely-connected Recurrent and Co-attentive Information

Open AccessPosted Content

Semantic Sentence Matching with Densely-connected Recurrent and Co-attentive Information

- 29 May 2018 -

TLDR

The authors proposed a densely-connected co-attentive recurrent neural network (C-RNN), which uses concatenated information of attentive features as well as hidden features of all the preceding recurrent layers.

Abstract:

Sentence matching is widely used in various natural language tasks such as natural language inference, paraphrase identification, and question answering. For these tasks, understanding logical and semantic relationship between two sentences is required but it is yet challenging. Although attention mechanism is useful to capture the semantic relationship and to properly align the elements of two sentences, previous methods of attention mechanism simply use a summation operation which does not retain original features enough. Inspired by DenseNet, a densely connected convolutional network, we propose a densely-connected co-attentive recurrent neural network, each layer of which uses concatenated information of attentive features as well as hidden features of all the preceding recurrent layers. It enables preserving the original and the co-attentive feature information from the bottommost word embedding layer to the uppermost recurrent layer. To alleviate the problem of an ever-increasing size of feature vectors due to dense concatenation operations, we also propose to use an autoencoder after dense concatenation. We evaluate our proposed architecture on highly competitive benchmark datasets related to sentence matching. Experimental results show that our architecture, which retains recurrent and attentive features, achieves state-of-the-art performances for most of the tasks.

Semantic Sentence Matching with Densely-connected Recurrent and Co-attentive Information

Citations

Text Feature Extraction and Selection Based on Attention Mechanism

Dynamic Feature Generation Network for Answer Selection

References

Deep Residual Learning for Image Recognition

Glove: Global Vectors for Word Representation

Densely Connected Convolutional Networks

Distributed Representations of Words and Phrases and their Compositionality

Distributed Representations of Words and Phrases and their Compositionality

Related Papers (5)

Glove: Global Vectors for Word Representation

A large annotated corpus for learning natural language inference

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Adam: A Method for Stochastic Optimization

Deep contextualized word representations