Temporal Difference Uncertainties as a Signal for Exploration

Open AccessJournal Article

Temporal Difference Uncertainties as a Signal for Exploration

Sebastian Flennerhag, +9 more

- 04 May 2021 -

arXiv: Artificial Intelligence

Chats0

TLDR

A novel method for estimating uncertainty over the value function that relies on inducing a distribution over temporal difference errors and incorporates exploration as an intrinsic reward and treats exploration as a separate learning problem, induced by the agent's temporal difference uncertainties.

Abstract:

An effective approach to exploration in reinforcement learning is to rely on an agent's uncertainty over the optimal policy, which can yield near-optimal exploration strategies in tabular settings. However, in non-tabular settings that involve function approximators, obtaining accurate uncertainty estimates is almost as challenging as the exploration problem itself. In this paper, we highlight that value estimates are easily biased and temporally inconsistent. In light of this, we propose a novel method for estimating uncertainty over the value function that relies on inducing a distribution over temporal difference errors. This exploration signal controls for state-action transitions so as to isolate uncertainty in value that is due to uncertainty over the agent's parameters. Because our measure of uncertainty conditions on state-action transitions, we cannot act on this measure directly. Instead, we incorporate it as an intrinsic reward and treat exploration as a separate learning problem, induced by the agent's temporal difference uncertainties. We introduce a distinct exploration policy that learns to collect data with high estimated uncertainty, which gives rise to a curriculum that smoothly changes throughout learning and vanishes in the limit of perfect value estimates. We evaluate our method on hard exploration tasks, including Deep Sea and Atari 2600 environments and find that our proposed form of exploration facilitates efficient exploration.

Temporal Difference Uncertainties as a Signal for Exploration

Citations

Semantic Exploration from Language Abstractions and Pretrained Representations

Sample Efficient Deep Reinforcement Learning via Uncertainty Estimation

Deciding What to Model: Value-Equivalent Sampling for Reinforcement Learning

Reinforcement Learning, Bit by Bit

Learning more skills through optimistic exploration.

References

Adam: A Method for Stochastic Optimization

Human-level control through deep reinforcement learning

Deep reinforcement learning with double Q-learning

On the likelihood that one unknown probability exceeds another in view of the evidence of two samples

Maximum entropy inverse reinforcement learning

Related Papers (5)

Smart exploration in reinforcement learning using absolute temporal difference errors

Model-Based Active Exploration

Learning and Using Models

Learning Temporal Point Processes via Reinforcement Learning

Robustness to Out-of-Distribution Inputs via Task-Aware Generative Uncertainty