Pregel: a system for large-scale graph processing

doi:10.1145/1807167.1807184

Proceedings ArticleDOI

Pregel: a system for large-scale graph processing

Grzegorz Malewicz, +6 more

- pp 135-146

Chats0

TLDR

A model for processing large graphs that has been designed for efficient, scalable and fault-tolerant implementation on clusters of thousands of commodity computers, and its implied synchronicity makes reasoning about programs easier.

Abstract:

Many practical computing problems concern large graphs. Standard examples include the Web graph and various social networks. The scale of these graphs - in some cases billions of vertices, trillions of edges - poses challenges to their efficient processing. In this paper we present a computational model suitable for this task. Programs are expressed as a sequence of iterations, in each of which a vertex can receive messages sent in the previous iteration, send messages to other vertices, and modify its own state and that of its outgoing edges or mutate graph topology. This vertex-centric approach is flexible enough to express a broad set of algorithms. The model has been designed for efficient, scalable and fault-tolerant implementation on clusters of thousands of commodity computers, and its implied synchronicity makes reasoning about programs easier. Distribution-related details are hidden behind an abstract API. The result is a framework for processing large graphs that is expressive and easy to program.

Citations

PDF

Open Access

More filters

Journal ArticleDOI

The future of scientific workflows

Ewa Deelman, +9 more

- 01 Jan 2018 -

International Journal of High Performanc...

TL;DR: This work highlights use cases, computing systems, workflow needs, and concludes by summarizing the remaining challenges this community sees that inhibit large-scale scientific workflows from becoming a mainstream tool for extreme-scale science.

...read moreread less

Proceedings ArticleDOI

NoDoze: Combatting Threat Alert Fatigue with Automated Provenance Triage

Wajih Ul Hassan, +6 more

TL;DR: NODOZE generates alert dependency graphs that are two orders of magnitude smaller than those generated by traditional tools without sacrificing the vital information needed for the investigation, and decreases the volume of false alarms by 84%, saving analysts’ more than 90 hours of investigation time per week.

...read moreread less

Journal ArticleDOI

Spinning fast iterative data flows

Stephan Ewen, +3 more

TL;DR: This work proposes a method to integrate incremental iterations, a form of workset iterations, with parallel dataflows and presents an extension to the programming model for incremental iterations that alleviates for the lack of mutable state in dataflow and allows for exploiting the sparse computational dependencies inherent in many iterative algorithms.

...read moreread less

Journal ArticleDOI

Scaling queries over big RDF graphs with semantic hash partitioning

Kisung Lee, +1 more

TL;DR: A novel semantic hash partitioning approach is presented and a Semantic HAsh Partitioning-Enabled distributed RDF data management system is implemented, called Shape, which scales well and can process big RDF datasets more efficiently than existing approaches.

...read moreread less

Proceedings ArticleDOI

GraphBIG: understanding graph computing in the context of industrial solutions

Lifeng Nai, +4 more

TL;DR: This paper characterized GraphBIG on real machines and observed extremely irregular memory patterns and significant diverse behavior across different computations, helping users understand the impact of modern graph computing on the hardware architecture and enables future architecture and system research.

...read moreread less

Collapse

References

PDF

Open Access

More filters

Journal ArticleDOI

A note on two problems in connexion with graphs

Edsger W. Dijkstra

- 01 Dec 1959 -

Numerische Mathematik

TL;DR: A tree is a graph with one and only one path between every two nodes, where at least one path exists between any two nodes and the length of each branch is given.

...read moreread less

Journal ArticleDOI

MapReduce: simplified data processing on large clusters

Jeffrey Dean, +1 more

TL;DR: This paper presents the implementation of MapReduce, a programming model and an associated implementation for processing and generating large data sets that runs on a large cluster of commodity machines and is highly scalable.

...read moreread less

Journal ArticleDOI

MapReduce: simplified data processing on large clusters

Jeffrey Dean, +1 more

- 01 Jan 2008 -

Communications of The ACM

TL;DR: This presentation explains how the underlying runtime system automatically parallelizes the computation across large-scale clusters of machines, handles machine failures, and schedules inter-machine communication to make efficient use of the network and disks.

...read moreread less

Journal ArticleDOI

The anatomy of a large-scale hypertextual Web search engine

Sergey Brin, +1 more

TL;DR: This paper provides an in-depth description of Google, a prototype of a large-scale search engine which makes heavy use of the structure present in hypertext and looks at the problem of how to effectively deal with uncontrolled hypertext collections where anyone can publish anything they want.

...read moreread less

Journal Article

The Anatomy of a Large-Scale Hypertextual Web Search Engine.

Sergey Brin, +1 more

- 01 Jan 1998 -

Computer Networks

TL;DR: Google as discussed by the authors is a prototype of a large-scale search engine which makes heavy use of the structure present in hypertext and is designed to crawl and index the Web efficiently and produce much more satisfying search results than existing systems.

...read moreread less

Collapse

Related Papers (5)

MapReduce: simplified data processing on large clusters

Jeffrey Dean, +1 more

- 01 Jan 2008 -

Communications of The ACM

GraphX: graph processing in a distributed dataflow framework

Joseph E. Gonzalez, +5 more

Pregel: a system for large-scale graph processing

Citations

The future of scientific workflows

NoDoze: Combatting Threat Alert Fatigue with Automated Provenance Triage

Spinning fast iterative data flows

Scaling queries over big RDF graphs with semantic hash partitioning

GraphBIG: understanding graph computing in the context of industrial solutions

References

A note on two problems in connexion with graphs

MapReduce: simplified data processing on large clusters

MapReduce: simplified data processing on large clusters

The anatomy of a large-scale hypertextual Web search engine

The Anatomy of a Large-Scale Hypertextual Web Search Engine.

Related Papers (5)

MapReduce: simplified data processing on large clusters

PowerGraph: distributed graph-parallel computation on natural graphs

Distributed GraphLab: a framework for machine learning and data mining in the cloud

A bridging model for parallel computation

GraphX: graph processing in a distributed dataflow framework