系统仿真学报 ›› 2026, Vol. 38 ›› Issue (7): 1964-1977.doi: 10.16182/j.issn1004731x.joss.25-0894

• 论文 • 上一篇    下一篇

基于分阶段引导的MADDPG多AUV围捕算法

张森1, 申四航1, 孙晓界1,2, 邵敬平1, 郭帅强1, 邓英杰3   

  1. 1.河南科技大学 信息工程学院,河南 洛阳 471000
    2.上海交通大学,上海 200240
    3.燕山大学,河北 秦皇岛 066004
  • 收稿日期:2025-09-15 修回日期:2025-12-22 出版日期:2026-07-28 发布日期:2026-07-31
  • 通讯作者: 孙晓界
  • 第一作者简介:张森(1984-),男,副教授,博士,研究方向为先进机器人与智能控制技术。
  • 基金资助:
    中国博士后科学基金面上项目(2025M770290);河南省自然科学基金(252300423316);河南省科技攻关项目(232102220016)

Multi-AUV Pursuit Algorithm with Phased Guidance Based on MADDPG

Zhang Sen1, Shen Sihang1, Sun Xiaojie1,2, Shao Jingping1, Guo Shuaiqiang1, Deng Yingjie3   

  1. 1.School of Information Engineering, Henan University of Science and Technology, Luoyang 471000, China
    2.Shanghai Jiao Tong University, Shanghai 200240, China
    3.Yanshan University, Qinhuangdao 066004, China
  • Received:2025-09-15 Revised:2025-12-22 Online:2026-07-28 Published:2026-07-31
  • Contact: Sun Xiaojie

摘要:

针对多AUV在基于MADDPG的目标围捕复杂环境下会出现训练初期探索和采样效率低下以及奖励稀疏等问题,提出了一种基于分阶段引导与课程学习的MADDPG(phased-guidance curriculum MADDPG,PGC-MADDPG)算法。通过课程学习将围捕任务分为目标跟踪与合围2个阶段;在目标跟踪阶段引入基于APF的经验策略作为引导项,提供目标方向的先验知识,加速AUV训练速度;在进入合围阶段后移除APF引导项,使AUV依靠自身策略完成围捕任务。对比实验结果表明:该算法与课程学习MADDPG算法相比,收敛速度提升了约36%,围捕成功率提高约5%。

关键词: 协同围捕, MADDPG, 课程学习, APF, AUV

Abstract:

To address the problems such as low exploration and sampling efficiency and sparse rewards in the early training stage under the complex target pursuit environment of multiple AUVs based on MADDPG, a phased-guidance curriculum MADDPG (PGC-MADDPG) algorithm was proposed. The pursuit task was divided into two phases, i.e., target tracking and encircling, through curriculum learning. In the target tracking phase, an experience strategy based on the APF method was introduced as a guidance item to provide prior knowledge of the target direction and accelerate the AUV training speed. After entering the encircling phase, the APF guidance item was removed, enabling the AUVs to accomplish the pursuit task relying on their own strategies. Comparative experimental results indicate that compared with the curriculum learning MADDPG algorithm, the proposed algorithm improves the convergence speed by approximately 36% and increases the final pursuit success rate by approximately 5%.

Key words: cooperative pursuit, MADDPG, curriculum learning, APF, AUV

中图分类号: