Reinforcement Learning: An Introduction

Open AccessBook

Reinforcement Learning: An Introduction

Chats0

TLDR

This book provides a clear and simple account of the key ideas and algorithms of reinforcement learning, which ranges from the history of the field's intellectual foundations to the most recent developments and applications.

Abstract:

Reinforcement learning, one of the most active research areas in artificial intelligence, is a computational approach to learning whereby an agent tries to maximize the total amount of reward it receives when interacting with a complex, uncertain environment. In Reinforcement Learning, Richard Sutton and Andrew Barto provide a clear and simple account of the key ideas and algorithms of reinforcement learning. Their discussion ranges from the history of the field's intellectual foundations to the most recent developments and applications. The only necessary mathematical background is familiarity with elementary concepts of probability. The book is divided into three parts. Part I defines the reinforcement learning problem in terms of Markov decision processes. Part II provides basic solution methods: dynamic programming, Monte Carlo methods, and temporal-difference learning. Part III presents a unified view of the solution methods and incorporates artificial neural networks, eligibility traces, and planning; the two final chapters present case studies and consider the future of reinforcement learning.

Citations

PDF

Open Access

More filters

Book

Deep Learning

Ian Goodfellow, +2 more

TL;DR: Deep learning as mentioned in this paper is a form of machine learning that enables computers to learn from experience and understand the world in terms of a hierarchy of concepts, and it is used in many applications such as natural language processing, speech recognition, computer vision, online recommendation systems, bioinformatics, and videogames.

...read moreread less

Journal ArticleDOI

Deep learning in neural networks

Jürgen Schmidhuber

- 01 Jan 2015 -

Neural Networks

TL;DR: This historical survey compactly summarizes relevant work, much of it from the previous millennium, review deep supervised learning, unsupervised learning, reinforcement learning & evolutionary computation, and indirect search for short programs encoding deep and large networks.

...read moreread less

Pattern Recognition and Machine Learning

Christopher M. Bishop

TL;DR: Probability distributions of linear models for regression and classification are given in this article, along with a discussion of combining models and combining models in the context of machine learning and classification.

...read moreread less

Collapse

References

PDF

Open Access

More filters

Journal ArticleDOI

Convergence Results for Single-Step On-PolicyReinforcement-Learning Algorithms

Satinder Singh, +3 more

- 01 Mar 2000 -

Machine Learning

TL;DR: This paper examines the convergence of single-step on-policy RL algorithms for control with both decaying exploration and persistent exploration and provides examples of exploration strategies that result in convergence to both optimal values and optimal policies.

...read moreread less

Journal ArticleDOI

Sample mean based index policies by O(log n) regret for the multi-armed bandit problem

Rajeev Agrawal

- 01 Dec 1995 -

Advances in Applied Probability

TL;DR: This paper constructs index policies that depend on the rewards from each arm only through their sample mean, and achieves a O(log n) regret with a constant that is based on the Kullback–Leibler number.

...read moreread less

Journal ArticleDOI

What is Intrinsic Motivation? A Typology of Computational Approaches.

Pierre-Yves Oudeyer, +1 more

- 02 Nov 2007 -

Frontiers in Neurorobotics

TL;DR: This paper sets the ground for a systematic operational study of intrinsic motivation by presenting a formal typology of possible computational approaches, partly based on existing computational models, but also presents new ways of conceptualizing intrinsic motivation.

...read moreread less

Book

Adaptive behavior and learning

John Staddon

TL;DR: The evolution, development and modification of behaviour, including feeding regulation, and Variation and selection of behaviour are studied.

...read moreread less

A Model of How the Basal Ganglia Generate and Use Neural Signals That Predict Reinforcement

James C. Houk, +2 more

TL;DR: This chapter contains sections titled: Introduction, Dopamine Neurons, Organization of Strtosomal Modules, Mechanism of Responsiveness to Predictors of Reinforcement, Correspondence with the Theory of Adaptive Critics, and Relation to the Actor-Critic Architecture.

...read moreread less