Journal Article
UA-FusionDet: Unregistered-Aware Infrared-Visible Fusion with Cross-Modal Feature Alignment for Maritime Ship Perception
Runbang Liu; Zhiyu Zhu; Huilin Ge; Jing Wang; Yongdong Shu; Qingshan Ji
Journal of Marine Science and Engineering · Vol. 14, Issue 17 · pp. 1611 · 2026
Abstract
Infrared and visible images provide complementary cues for maritime ship detection, but practical dual-sensor systems often produce image pairs that are not strictly registered. Directly fusing such unregistered pairs may introduce ghosting artifacts, blurred target boundaries, and feature conflicts, especially over weak-texture sea surfaces where reliable correspondence cues are sparse. To address this problem, we propose UA-FusionDet, an unregistered-aware infrared-visible fusion framework with cross-modal feature alignment for maritime ship detection. The proposed framework extracts visible and infrared features with a dual-branch encoder, aligns the visible feature to the infrared reference through cross-modal deformable feature alignment, and suppresses unstable background offsets using sea-surface saliency guidance. Wavelet-guided complementary fusion then decomposes the aligned features into low- and high-frequency sub-bands, enabling frequency-aware fusion of infrared thermal saliency and visible structural details before feeding the fused representation to both a lightweight reconstruction decoder and a ship detection head. The reconstruction decoder provides auxiliary image-level regularization, while the detection branch supervises the task-oriented fused representation with ship bounding-box annotations. UA-FusionDet does not require registration ground truth or fused-image ground truth during training, making it suitable for realistic maritime monitoring scenarios with imperfectly aligned visible and infrared sensors. Experiments on 3132 unregistered visible–LWIR maritime image pairs show that UA-FusionDet achieves a precision of 0.904, a recall of 0.866, an mAP50 of 0.912, and an mAP50–95 of 0.566, exceeding the strongest competing method by 2.5 and 2.4 percentage points on the two mAP metrics, respectively, while maintaining an inference speed of 52.6 FPS. These results demonstrate that the proposed alignment and fusion framework improves detection accuracy under cross-modal misregistration while retaining practical inference efficiency.