AIBPO: Combine the Intrinsic Reward and Auxiliary Task for 3D Strategy Game

doi:10.1155/2021/6698231

Open AccessJournal ArticleDOI

AIBPO: Combine the Intrinsic Reward and Auxiliary Task for 3D Strategy Game

Huale Li, +7 more

- 13 Jul 2021 -

Complexity

- Vol. 2021, pp 1-9

TLDR

Zhang et al. as mentioned in this paper proposed an intrinsic-based policy optimization (IBPO) algorithm for reward sparsity, where a novel intrinsic reward is integrated into the value network, which provides an additional reward in the environment with sparse reward, so as to accelerate the training.

Abstract:

In recent years, deep reinforcement learning (DRL) achieves great success in many fields, especially in the field of games, such as AlphaGo, AlphaZero, and AlphaStar. However, due to the reward sparsity problem, the traditional DRL-based method shows limited performance in 3D games, which contain much higher dimension of state space. To solve this problem, in this paper, we propose an intrinsic-based policy optimization (IBPO) algorithm for reward sparsity. In the IBPO, a novel intrinsic reward is integrated into the value network, which provides an additional reward in the environment with sparse reward, so as to accelerate the training. Besides, to deal with the problem of value estimation bias, we further design three types of auxiliary tasks, which can evaluate the state value and the action more accurately in 3D scenes. Finally, a framework of auxiliary intrinsic-based policy optimization (AIBPO) is proposed, which improves the performance of the IBPO. The experimental results show that the method is able to deal with the reward sparsity problem effectively. Therefore, the proposed method may be applied to real-world scenarios, such as 3-dimensional navigation and automatic driving, which can improve the sample utilization to reduce the cost of interactive sample collected by the real equipment.

AIBPO: Combine the Intrinsic Reward and Auxiliary Task for 3D Strategy Game

Citations

Heterogeneous catalysis mediated by light, electricity and enzyme via machine learning: Paradigms, applications and prospects.

References

Deep learning

Human-level control through deep reinforcement learning

Mastering the game of Go with deep neural networks and tree search

Mastering the game of Go without human knowledge

Learning to Forget: Continual Prediction with LSTM

Related Papers (5)

Towards Designing Optimal Reward Functions in Multi-Agent Reinforcement Learning Problems

Self-Supervised Online Reward Shaping in Sparse-Reward Environments.

Analyzing and visualizing multiagent rewards in dynamic and stochastic domains

Keeping Your Distance: Solving Sparse Reward Tasks Using Self-Balancing Shaped Rewards

Toward Diverse Text Generation with Inverse Reinforcement Learning