Abstract
The safe and efficient collision avoidance of multiple ships is essential for maritime navigation and intelligent shipping systems. In this paper, we propose a novel COLREGs-compliant multi-ship collision avoidance strategy based on deep reinforcement learning. A cooperative training framework using the Proximal Policy Optimization (PPO) algorithm enables multiple ship agents to learn optimal collision avoidance actions while considering the interactions and motions of neighboring ships. Encounter situation awareness mechanisms and carefully designed reward functions are integrated to ensure strict adherence to the International Regulations for Preventing Collisions at Sea (COLREGs), while a multi-objective optimization approach embedded in the reward function balances collision risk, navigational efficiency, route smoothness, and destination achievement. Extensive simulations covering diverse ship encounter scenarios demonstrate the effectiveness, robustness, and COLREGs compliance of the proposed strategy, highlighting its practical potential for multi-ship navigation systems.