Abstract
In this study, we propose the Frequency-domain Feature Fusion Module (F3M) to address the challenges of underwater object detection, where optical degradation—particularly high-frequency attenuation and low-frequency color distortion—significantly compromises performance. We critically re-evaluate the need for strict invertibility in detection-oriented frequency modeling. Traditional wavelet-based methods incur high computational redundancy to maintain signal reconstruction, whereas F3M introduces a lightweight “Separate–Project–Fuse” paradigm. This mechanism decouples low-frequency illumination artifacts from high-frequency structural cues via spatial approximation, enabling the recovery of fine-scale details like coral textures and debris boundaries without the overhead of channel expansion. We validate F3M’s versatility by integrating it into both Convolutional Neural Networks (YOLO) and Transformer-based detectors (RT-DETR). Evaluations on the SCoralDet dataset show consistent improvements: F3M enhances the lightweight YOLO11n by 3.5% mAP50 and increases RT-DETR-n’s localization accuracy (mAP50–95) from 0.514 to 0.532. Additionally, cross-domain validation on the deep-sea TrashCan-Instance dataset shows F3M achieving comparable accuracy to the larger YOLOv8n while requiring 13% fewer parameters and 20% fewer GFLOPs. This study confirms that frequency-domain modulation provides an efficient and widely applicable enhancement for real-time underwater perception.