Mastering the game of Go with deep neural networks and tree search

doi:10.1038/NATURE16961

Journal ArticleDOI

Mastering the game of Go with deep neural networks and tree search

David Silver, +19 more

- 28 Jan 2016 -

Nature

- Vol. 529, Iss: 7587, pp 484-489

TLDR

Using this search algorithm, the program AlphaGo achieved a 99.8% winning rate against other Go programs, and defeated the human European Go champion by 5 games to 0.5, the first time that a computer program has defeated a human professional player in the full-sized game of Go.

Abstract:

The game of Go has long been viewed as the most challenging of classic games for artificial intelligence owing to its enormous search space and the difficulty of evaluating board positions and moves. Here we introduce a new approach to computer Go that uses ‘value networks’ to evaluate board positions and ‘policy networks’ to select moves. These deep neural networks are trained by a novel combination of supervised learning from human expert games, and reinforcement learning from games of self-play. Without any lookahead search, the neural networks play Go at the level of stateof-the-art Monte Carlo tree search programs that simulate thousands of random games of self-play. We also introduce a new search algorithm that combines Monte Carlo simulation with value and policy networks. Using this search algorithm, our program AlphaGo achieved a 99.8% winning rate against other Go programs, and defeated the human European Go champion by 5 games to 0. This is the first time that a computer program has defeated a human professional player in the full-sized game of Go, a feat previously thought to be at least a decade away.

Citations

PDF

Open Access

More filters

Journal ArticleDOI

Dermatologist-level classification of skin cancer with deep neural networks

Andre Esteva, +7 more

- 02 Feb 2017 -

Nature

TL;DR: This work demonstrates an artificial intelligence capable of classifying skin cancer with a level of competence comparable to dermatologists, trained end-to-end from images directly, using only pixels and disease labels as inputs.

...read moreread less

Journal ArticleDOI

Mastering the game of Go without human knowledge

David Silver, +16 more

- 19 Oct 2017 -

Nature

TL;DR: An algorithm based solely on reinforcement learning is introduced, without human data, guidance or domain knowledge beyond game rules, that achieves superhuman performance, winning 100–0 against the previously published, champion-defeating AlphaGo.

...read moreread less

Proceedings ArticleDOI

Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization

Ramprasaath R. Selvaraju, +5 more

TL;DR: This work combines existing fine-grained visualizations to create a high-resolution class-discriminative visualization, Guided Grad-CAM, and applies it to image classification, image captioning, and visual question answering (VQA) models, including ResNet-based architectures.

...read moreread less

Proceedings ArticleDOI

Towards Evaluating the Robustness of Neural Networks

Nicholas Carlini, +1 more

TL;DR: In this paper, the authors demonstrate that defensive distillation does not significantly increase the robustness of neural networks by introducing three new attack algorithms that are successful on both distilled and undistilled neural networks with 100% probability.

...read moreread less

Journal ArticleDOI

Places: A 10 Million Image Database for Scene Recognition

Bolei Zhou, +4 more

- 01 Jun 2018 -

IEEE Transactions on Pattern Analysis an...

TL;DR: The Places Database is described, a repository of 10 million scene photographs, labeled with scene semantic categories, comprising a large and diverse list of the types of environments encountered in the world, using the state-of-the-art Convolutional Neural Networks as baselines, that significantly outperform the previous approaches.

...read moreread less

Collapse

References

PDF

Open Access

More filters

Journal Article

From simple features to sophisticated evaluation functions

Michael Buro

- 01 Jan 1999 -

Lecture Notes in Computer Science

TL;DR: A practical framework for the semi-automatic construction of evaluation-functions for games based on a structured evaluation function representation is presented that is able to discover new features in a computationally feasible way.

...read moreread less

Proceedings ArticleDOI

Monte-Carlo simulation balancing

David Silver, +1 more

TL;DR: The main idea is to optimise the balance of a simulation policy, so that an accurate spread of simulation outcomes is maintained, rather than optimising the direct strength of the simulation policy.

...read moreread less

Proceedings Article

Temporal difference learning applied to a high-performance game-playing program

Jonathan Schaeffer, +2 more

TL;DR: This paper shows that TD learinng is capable of competing with the best human effort.

...read moreread less

Book ChapterDOI

Whole-History Rating: A Bayesian Rating System for Players of Time-Varying Strength

Rémi Coulom

TL;DR: Experiments demonstrate that, in comparison to Elo, Glicko, TrueSkill, and decayed-history algorithms, WHR produces better predictions.

...read moreread less

Proceedings ArticleDOI

Bayesian pattern ranking for move prediction in the game of Go

David L. Stern, +2 more

TL;DR: A probability distribution over legal moves for professional play in a given position in Go is obtained and shows excellent prediction performance as indicated by its ability to perfectly predict the moves made by professional Go players in 34% of test positions.

...read moreread less