Relative Entropy of Correct Proximal Policy Optimization Algorithms with Modified Penalty Factor in Complex Environment

Weimin Chen; Kelvin Kian Loong Wong; Sifan Long; Zhili Sun

doi:10.3390/e24040440

Back

Relative Entropy of Correct Proximal Policy Optimization Algorithms with Modified Penalty Factor in Complex Environment

Journal article

Open access

Peer reviewed

Relative Entropy of Correct Proximal Policy Optimization Algorithms with Modified Penalty Factor in Complex Environment

Weimin Chen, Kelvin Kian Loong Wong, Sifan Long and Zhili Sun

Entropy (Basel, Switzerland), Vol.24(4), p.440

22/03/2022

DOI: https://doi.org/10.3390/e24040440

PMID: 35455103

Abstract

approximation theory

correct proximal policy optimization

entropy

policy gradient

reinforcement learning

Optimization

In the field of reinforcement learning, we propose a Correct Proximal Policy Optimization (CPPO) algorithm based on the modified penalty factor and relative entropy in order to solve the robustness and stationarity of traditional algorithms. Firstly, In the process of reinforcement learning, this paper establishes a strategy evaluation mechanism through the policy distribution function. Secondly, the state space function is quantified by introducing entropy, whereby the approximation policy is used to approximate the real policy distribution, and the kernel function estimation and calculation of relative entropy is used to fit the reward function based on complex problem. Finally, through the comparative analysis on the classic test cases, we demonstrated that our proposed algorithm is effective, has a faster convergence speed and better performance than the traditional PPO algorithm, and the measure of the relative entropy can show the differences. In addition, it can more efficiently use the information of complex environment to learn policies. At the same time, not only can our paper explain the rationality of the policy distribution theory, the proposed framework can also balance between iteration steps, computational complexity and convergence speed, and we also introduced an effective measure of performance using the relative entropy concept.

Files and links (1)

url

https://doi.org/10.3390/e24040440View

Published (Version of record) Open

Metrics

16 Record Views

3 Times Cited - Web of Science

Details

Title: Relative Entropy of Correct Proximal Policy Optimization Algorithms with Modified Penalty Factor in Complex Environment
Creators: Weimin Chen - Hunan City University
Kelvin Kian Loong Wong - Hunan City University
Sifan Long - Central South University
Zhili Sun - University of Surrey
Publication Details: Entropy (Basel, Switzerland), Vol.24(4), p.440
Date published: 22/03/2022
Identifiers: 99817242902346
Academic Unit: School of Computer Science and Electronic Engineering
Language: English
Resource Type: Journal article

Relative Entropy of Correct Proximal Policy Optimization Algorithms with Modified Penalty Factor in Complex Environment

Abstract

Files and links (1)

Metrics

Details

Usage Policy