Streaming Similarity Search over One Billion Tweets Using Parallel Locality-Sensitive Hashing

Open Access

Streaming Similarity Search over One Billion Tweets Using Parallel Locality-Sensitive Hashing

Chats0

TLDR

Parallel LSH (PLSH) as mentioned in this paper is a variant of LSH designed to be extremely efficient, capable of scaling out on multiple nodes and multiple cores, and which supports high-throughput streaming of new data.

Abstract:

Finding nearest neighbors has become an important operation on databases, with applications to text search, multimedia indexing, and many other areas. One popular algorithm for similarity search, especially for high dimensional data (where spatial indexes like kd-trees do not perform well) is Locality Sensitive Hashing (LSH), an approximation algorithm for finding similar objects.In this paper, we describe a new variant of LSH, called Parallel LSH (PLSH) designed to be extremely efficient, capable of scaling out on multiple nodes and multiple cores, and which supports high-throughput streaming of new data. Our approach employs several novel ideas, including: cache-conscious hash table layout, using a 2-level merge algorithm for hash table construction; an efficient algorithm for duplicate elimination during hash-table querying; an insert-optimized hash table structure and efficient data expiration algorithm for streaming data; and a performance model that accurately estimates performance of the algorithm and can be used to optimize parameter settings. We show that on a workload where we perform similarity search on a dataset of > 1 Billion tweets, with hundreds of millions of new tweets per day, we can achieve query times of 1-2.5 ms. We show that this is an order of magnitude faster than existing indexing schemes, such as inverted indexes. To the best of our knowledge, this is the fastest implementation of LSH, with table construction times up to 3.7× faster and query times that are 8.3× faster than a basic implementation.

Streaming Similarity Search over One Billion Tweets Using Parallel Locality-Sensitive Hashing

Citations

Approximate Nearest Neighbor Search in High Dimensions

FLASH: Randomized Algorithms Accelerated over CPU-GPU for Ultra-High Dimensional Similarity Search

The Unreasonable Effectiveness of Structured Random Orthogonal Embeddings

Scalability and Total Recall with Fast CoveringLSH

Massively-Parallel Similarity Join, Edge-Isoperimetry, and Distance Correlations on the Hypercube

References

Multidimensional binary search trees used for associative searching

Approximate nearest neighbors: towards removing the curse of dimensionality

Direct methods for sparse matrices

Near-optimal hashing algorithms for approximate nearest neighbor in high dimensions

Robust and fast similarity search for moving object trajectories

Related Papers (5)

Streaming similarity search over one billion tweets using parallel locality-sensitive hashing

Frequency Based Locality Sensitive Hashing

A Generic Method for Accelerating LSH-Based Similarity Join Processing

Similarity estimation techniques from rounding algorithms

An Efficient Similarity Searching Scheme in Massive Databases