Abstract
Reliable extraction of dolphin whistle contours is fundamental for scalable analyses of vocal identity, repertoire structure, and individual‐based acoustic monitoring. We evaluated four algorithms for fundamental frequency estimation of common bottlenose dolphin ( Tursiops truncatus ) whistles, leveraging a benchmark dataset of annotated whistles from known animals to assess performance across individuals and as a function of signal‐to‐noise ratio (SNR). In high‐SNR whistles (> 20 dB), the deep learning models CREPE‐tt, SAM‐whistle, and Silbido Profundo achieved similar performance in mean contour coverage (~80%), with CREPE‐tt producing the most continuous contours and the lowest frequency error. SAM‐whistle and Silbido Profundo were more robust than CREPE‐tt under low‐SNR conditions. However, automatically extracted contours consistently reduced within‐ versus between‐individual separation and ultimately decreased individual classification accuracy, highlighting how contour gaps and fragmentation propagate into downstream identity matching. Taken together, our results suggest a practical division: CREPE‐tt excels with high‐SNR data and detailed contour shape analyses without harmonic post‐processing, whereas SAM‐whistle and Silbido Profundo perform better with low‐SNR recordings. For individual identification and abundance estimation, explicit SNR‐based inclusion thresholds and quality‐control criteria remain necessary.