生物多样性

• • 上一篇    下一篇

声纹与空间分布联合的中国鸟类鸣声识别方法研究

孔冕, 李春明, 万世龙, 崔胜辉   

  1. 中国科学院城市环境研究所区域与城市生态安全全国重点实验室, 福建 361021 中国
    中国科学院大学, 北京 100049 中国
  • 收稿日期:2026-02-28 修回日期:2026-04-23
  • 通讯作者: 李春明
  • 基金资助:
    国家自然科学基金项目(42277472); “十四五”国家重点研发子课题(2023YFF1304600)

A Joint Acoustic and Spatial Distribution Approach for Bird Vocalization Recognition in China

Mian Kong, Chunming Li, Shilong Wan, Shenghui Cui   

  1. State Key Laboratory of Regional and Urban Ecology, Institute of Urban Environment, Chinese Academy of Sciences 361021, China
    , University of Chinese Academy of Sciences 100049, China
  • Received:2026-02-28 Revised:2026-04-23
  • Contact: Chunming Li
  • Supported by:
    National Natural Science Foundation of China(42277472); National Key R&D Program of China(2023YFF1304600)

摘要: 鸟类是生物多样性保护的关键指示类群,能敏感响应栖息地变化。鸣声识别模型数据集的不完整,物种地理分布信息的未广泛整合,以及与实地应用对比验证的缺乏,制约了我国鸟类鸣声识别模型的发展。本研究构建了一个包含中国1368种鸟类鸣声的数据集,建立了基于迁移学习的两阶段训练的鸟类鸣声声纹识别模型,以及利用公民观鸟数据构建的涵盖1428种鸟类的市级尺度的空间分布概率模型,最后通过Sigmoid函数建立了声纹与空间分布联合的鸟类鸣声识别模型并开展应用。在模型性能层面,选取了66种全国常见鸟类鸣声数据集,开展了模型性能验证。发现随着置信度的提高,联合模型比声纹识别模型表现更稳健,准确度最高提升了16.8%。在实际应用能力层面,本研究基于4年站点监测数据和人工同步观测数据进行对比分析,在物种数量方面机器共识别到68种,占人工结果的60.2%,机器识别更能准确反映鸟类活动周期,并能更快速的响应鸟类的活动规律。

关键词: 鸟类, 声纹识别, 中国鸟类数据库, 空间分布模型, 联合模型

AbstractBirds are key indicator taxa for biodiversity conservation due to their sensitive responses to habitat changes. The development of bird vocalization recognition models in China is constrained by incomplete datasets for model training, insufficient integration of geographical distribution information, and a lack of field validation through comparative studies. To address these gaps, this study constructed an acoustic dataset comprising vocalizations of 1,368 bird species in China. A bird vocalization recognition model was developed using a two-stage transfer learning approach. Concurrently, a species spatial distribution probability model at the municipal scale, covering 1,428 species, was established based on citizen science bird observation data. Finally, a joint bird vocalization recognition model integrating acoustic and spatial distribution information via a Sigmoid function was developed and applied. Model performance was validated using a dataset of vocalizations from 66 common nationwide bird species. The results indicate that as the confidence threshold increases, the joint model demonstrates more stable performance than the standalone acoustic recognition model, with a maximum accuracy improvement of 16.8%. For practical application, a comparative analysis is conducted using four years of data from site-based monitoring and manual synchronous observation. In terms of species richness, the machine identifies 68 species, accounting for 60.2% of the species recorded manually. The machine-based method reflects the activity patterns of bird species accurately and responds more rapidly to changes in bird activity rhythms.

Key words: Bird vocalization, acoustic recognition, China bird database, species distribution model, joint model