Scalable Probabilistic Similarity Ranking in Uncertain Databases (Technical Report)

Open AccessPosted Content

Scalable Probabilistic Similarity Ranking in Uncertain Databases (Technical Report)

Thomas Bernecker, +4 more

- 16 Jul 2009 -

arXiv: Databases

Chats0

TLDR

A scalable approach for probabilistic top-k similarity ranking on uncertain vector data that reduces this to a linear-time complexity while having the same memory requirements, facilitated by incremental accessing of the uncertain vector instances in increasing order of their distance to a reference object.

Abstract:

This paper introduces a scalable approach for probabilistic top-k similarity ranking on uncertain vector data. Each uncertain object is represented by a set of vector instances that are assumed to be mutually-exclusive. The objective is to rank the uncertain data according to their distance to a reference object. We propose a framework that incrementally computes for each object instance and ranking position, the probability of the object falling at that ranking position. The resulting rank probability distribution can serve as input for several state-of-the-art probabilistic ranking models. Existing approaches compute this probability distribution by applying a dynamic programming approach of quadratic complexity. In this paper we theoretically as well as experimentally show that our framework reduces this to a linear-time complexity while having the same memory requirements, facilitated by incremental accessing of the uncertain vector instances in increasing order of their distance to the reference object. Furthermore, we show how the output of our method can be used to apply probabilistic top-k ranking for the objects, according to different state-of-the-art definitions. We conduct an experimental evaluation on synthetic and real data, which demonstrates the efficiency of our approach.

References

PDF

Open Access

More filters

Journal ArticleDOI

Efficient query evaluation on probabilistic databases

Nilesh Dalvi, +1 more

TL;DR: It is shown that the data complexity of some queries is #P-complete, which implies that these queries do not admit any efficient evaluation methods, and an optimization algorithm is described that can compute efficiently most queries.

...read moreread less

Proceedings ArticleDOI

Evaluating probabilistic queries over imprecise data

Reynold Cheng, +2 more

TL;DR: This paper addresses the important issue of measuring the quality of the answers to query evaluation based upon uncertain data, and provides algorithms for efficiently pulling data from relevant sensors or moving objects in order to improve thequality of the executing queries.

...read moreread less

Proceedings ArticleDOI

ULDBs: databases with uncertainty and lineage

Omar Benjelloun, +3 more

TL;DR: It is shown that the ULDB representation is complete, and that it permits straightforward implementation of many relational operations, and how ULDBs enable a new approach to query processing in probabilistic databases.

...read moreread less

Proceedings ArticleDOI

Top-k Query Processing in Uncertain Databases

Mohamed A. Soliman, +2 more

TL;DR: A framework that encapsulates a state space model and efficient query processing techniques to tackle the challenges of uncertain data settings is constructed and it is proved that the techniques are optimal in terms of the number of accessed tuples and materialized search states.

...read moreread less

Journal ArticleDOI

Updating and Querying Databases that Track Mobile Units

Ouri Wolfson, +3 more

TL;DR: The update problem is to determine when the location of a moving object in the database (namely its database location) should be updated, and an information cost model is proposed that captures uncertainty, deviation, and communication.

...read moreread less