Abstract
To address the control coupling challenges arising from task heterogeneity of unmanned surface vehicle (USV) formation, a distributed hybrid deep reinforcement learning (HDRL) framework is proposed. The framework decomposes the formation task into two subtasks: leader path planning using the single-agent deep reinforcement learning (SADRL) algorithm and follower formation tracking using the multi-agent deep reinforcement learning (MADRL) algorithm. By embedding the physical constraints of the real Otter USV into the training loop, the policy network outputs are mapped to propeller revolutions that conform to its dynamic characteristics. To optimize control performance, a dynamic gating mechanism triggered by formation position error is developed to mitigate multi-objective interference through temporal task scheduling. Concurrently, a mirror mapping mechanism leveraging the physical symmetry of the formation is designed to achieve policy sharing and data augmentation. Furthermore, the desired velocity calculated based on rigid-body kinematics is used to achieve kinematic-compensated formation tracking. The simulation results indicate that, compared to the SADRL algorithm, the planning success rate of HDRL is improved by 44.59%. Furthermore, compared to the MADRL algorithm, the integrated tracking performance is enhanced by 21.79–39.64%.