scispace - formally typeset
Open AccessProceedings ArticleDOI

Large-Scale Image Retrieval with Attentive Deep Local Features

Reads0
Chats0
TLDR
An attentive local feature descriptor suitable for large-scale image retrieval, referred to as DELE (DEep Local Feature), based on convolutional neural networks, which are trained only with image-level annotations on a landmark image dataset.
Abstract
We propose an attentive local feature descriptor suitable for large-scale image retrieval, referred to as DELE (DEep Local Feature). The new feature is based on convolutional neural networks, which are trained only with image-level annotations on a landmark image dataset. To identify semantically useful local features for image retrieval, we also propose an attention mechanism for key point selection, which shares most network layers with the descriptor. This frame-work can be used for image retrieval as a drop-in replacement for other keypoint detectors and descriptors, enabling more accurate feature matching and geometric verification. Our system produces reliable confidence scores to reject false positives–in particular, it is robust against queries that have no correct match in the database. To evaluate the proposed descriptor, we introduce a new large-scale dataset, referred to as Google-Landmarks dataset, which involves challenges in both database and query such as background clutter, partial occlusion, multiple landmarks, objects in variable scales, etc. We show that DELE outperforms the state-of-the-art global and local descriptors in the large-scale setting by significant margins.

read more

Citations
More filters
Proceedings ArticleDOI

Generalized Local Attention Pooling for Deep Metric Learning

TL;DR: In this paper, the authors proposed a generalized local attention pooling (GLAP) to generate compact image representations which uses local spatial information through an attention mechanism, instead of being placed at the end layer of the backbone, is connected at an intermediate level, resulting in lower memory requirements.
Book ChapterDOI

Content-Based Image Retrieval and the Semantic Gap in the Deep Learning Era

TL;DR: In this article, the authors show that the recent advances in instance retrieval transfer to more generic image retrieval scenarios, which is called instance or object retrieval and requires matching fine-grained visual patterns between images.
Proceedings ArticleDOI

Video Logo Retrieval Based on Local Features

TL;DR: An algorithm called Video Logo Retrieval (VLR), which is an image-to-video retrieval algorithm based on the spatial distribution of local image descriptors that measure the distance between the query image (the logo) and a collection of video images.
Journal ArticleDOI

Street-Level Image Localization Based on Building-Aware Features via Patch-Region Retrieval under Metropolitan-Scale

TL;DR: Zhang et al. as mentioned in this paper proposed a building-aware feature (BAF) and a patch-region retrieval method (PRR) for image-based localization under metropolitan scale.
Posted Content

Unsupervised Metric Relocalization Using Transform Consistency Loss

TL;DR: This work proposes a self-supervised solution to metric relocalization, which exploits a key insight: localizing a query image within a map should yield the same absolute pose, regardless of the reference image used for registration, and derives a novel transform consistency loss.
References
More filters
Proceedings ArticleDOI

Deep Residual Learning for Image Recognition

TL;DR: In this article, the authors proposed a residual learning framework to ease the training of networks that are substantially deeper than those used previously, which won the 1st place on the ILSVRC 2015 classification task.
Proceedings Article

Very Deep Convolutional Networks for Large-Scale Image Recognition

TL;DR: In this paper, the authors investigated the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting and showed that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 layers.
Journal ArticleDOI

Distinctive Image Features from Scale-Invariant Keypoints

TL;DR: This paper presents a method for extracting distinctive invariant features from images that can be used to perform reliable matching between different views of an object or scene and can robustly identify objects among clutter and occlusion while achieving near real-time performance.
Journal ArticleDOI

ImageNet Large Scale Visual Recognition Challenge

TL;DR: The ImageNet Large Scale Visual Recognition Challenge (ILSVRC) as mentioned in this paper is a benchmark in object category classification and detection on hundreds of object categories and millions of images, which has been run annually from 2010 to present, attracting participation from more than fifty institutions.
Journal ArticleDOI

Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography

TL;DR: New results are derived on the minimum number of landmarks needed to obtain a solution, and algorithms are presented for computing these minimum-landmark solutions in closed form that provide the basis for an automatic system that can solve the Location Determination Problem under difficult viewing.
Related Papers (5)