Journal of System Simulation ›› 2026, Vol. 38 ›› Issue (7): 1964-1977.doi: 10.16182/j.issn1004731x.joss.25-0894

• Papers • Previous Articles     Next Articles

Multi-AUV Pursuit Algorithm with Phased Guidance Based on MADDPG

Zhang Sen1, Shen Sihang1, Sun Xiaojie1,2, Shao Jingping1, Guo Shuaiqiang1, Deng Yingjie3   

  1. 1.School of Information Engineering, Henan University of Science and Technology, Luoyang 471000, China
    2.Shanghai Jiao Tong University, Shanghai 200240, China
    3.Yanshan University, Qinhuangdao 066004, China
  • Received:2025-09-15 Revised:2025-12-22 Online:2026-07-28 Published:2026-07-31
  • Contact: Sun Xiaojie

Abstract:

To address the problems such as low exploration and sampling efficiency and sparse rewards in the early training stage under the complex target pursuit environment of multiple AUVs based on MADDPG, a phased-guidance curriculum MADDPG (PGC-MADDPG) algorithm was proposed. The pursuit task was divided into two phases, i.e., target tracking and encircling, through curriculum learning. In the target tracking phase, an experience strategy based on the APF method was introduced as a guidance item to provide prior knowledge of the target direction and accelerate the AUV training speed. After entering the encircling phase, the APF guidance item was removed, enabling the AUVs to accomplish the pursuit task relying on their own strategies. Comparative experimental results indicate that compared with the curriculum learning MADDPG algorithm, the proposed algorithm improves the convergence speed by approximately 36% and increases the final pursuit success rate by approximately 5%.

Key words: cooperative pursuit, MADDPG, curriculum learning, APF, AUV

CLC Number: