The ground truth about metadata and community detection in networks

doi:10.1126/SCIADV.1602548

Open AccessJournal ArticleDOI

The ground truth about metadata and community detection in networks

Leto Peel, +4 more

- 01 May 2017 -

Science Advances

- Vol. 3, Iss: 5, pp 1602548

Chats0

TLDR

It is proved that no algorithm can uniquely solve community detection, and a general No Free Lunch theorem for community detection is proved, which implies that there can be no algorithm that is optimal for all possible community detection tasks.

Abstract:

Across many scientific domains, there is a common need to automatically extract a simplified view or coarse-graining of how a complex system's components interact. This general task is called community detection in networks and is analogous to searching for clusters in independent vector data. It is common to evaluate the performance of community detection algorithms by their ability to find so-called ground truth communities. This works well in synthetic networks with planted communities because these networks' links are formed explicitly based on those known communities. However, there are no planted communities in real-world networks. Instead, it is standard practice to treat some observed discrete-valued node attributes, or metadata, as ground truth. We show that metadata are not the same as ground truth and that treating them as such induces severe theoretical and practical problems. We prove that no algorithm can uniquely solve community detection, and we prove a general No Free Lunch theorem for community detection, which implies that there can be no algorithm that is optimal for all possible community detection tasks. However, community detection remains a powerful tool and node metadata still have value, so a careful exploration of their relationship with network structure can yield insights of genuine worth. We illustrate this point by introducing two statistical techniques that can quantify the relationship between metadata and community structure for a broad class of models. We demonstrate these techniques using both synthetic and real-world networks, and for multiple types of metadata and community structures.

The ground truth about metadata and community detection in networks

Citations

Community Discovery in Dynamic Networks: A Survey

Global energy flows embodied in international trade: A combination of environmentally extended input–output analysis and complex network analysis

Metrics for Community Analysis: A Survey

The modular organization of human anatomical brain networks: Accounting for the cost of wiring

Community detection in node-attributed social networks: A survey

References

Link communities reveal multiscale complexity in networks

Defining and evaluating network communities based on ground-truth

The lack of a priori distinctions between learning algorithms

Modern Multidimensional Scaling: Theory and Applications

Estimation and prediction for stochastic blockstructures

Related Papers (5)

Fast unfolding of communities in large networks

Finding and evaluating community structure in networks.

Community detection in graphs

Community structure in social and biological networks

Modularity and community structure in networks