系统工程与电子技术 ›› 2026, Vol. 48 ›› Issue (9): 2898-2906.doi: 10.12305/j.issn.1001-506X.2026.09.02

• 电子技术 • 上一篇    

基于混合扩散模型的水声数据增扩方法

马治勋1,2(), 汤宁1(), 李璇1,2(), 郝程鹏1,2()   

  1. 1. 中国科学院声学研究所,北京 100190
    2. 中国科学院大学电子电气与通信工程学院,北京 100049
  • 收稿日期:2025-09-11 修回日期:2025-11-10 接受日期:2025-12-02 出版日期:2026-01-20 发布日期:2026-01-20
  • 通讯作者: 郝程鹏 E-mail:mazhixun@mail.ioa.ac.cn;tangning@mail.ioa.ac.cn;lixuan@mail.ioa.ac.cn;haochengp@mail.ioa.ac.cn
  • 作者简介:马治勋(1991—),男,助理研究员,硕士,主要研究方向为水声信号处理、自适应目标检测与目标识别
    汤 宁(1994—),女,助理研究员,博士,主要研究方向为水声信号处理
    李 璇(1983—),女,研究员,博士,主要研究方向为水声信号处理、目标识别

Mixed diffusion model-based underwater acoustic data augmentation method

Zhixun Ma1,2(), Ning Tang1(), Xuan Li1,2(), Chengpeng Hao1,2()   

  1. 1. Institute of Acoustics,Chinese Academy of Sciences,Beijing 100190,China
    2. School of Electronic,Electrical and Communication Engineering,University of Chinese Academy of Sciences,Beijing 100049,China
  • Received:2025-09-11 Revised:2025-11-10 Accepted:2025-12-02 Online:2026-01-20 Published:2026-01-20
  • Contact: Chengpeng Hao E-mail:mazhixun@mail.ioa.ac.cn;tangning@mail.ioa.ac.cn;lixuan@mail.ioa.ac.cn;haochengp@mail.ioa.ac.cn

摘要:

针对水声目标识别领域广泛存在的数据稀缺性问题,提出一种基于混合扩散模型的数据增扩方法。为区别于传统频谱变换或生成对抗网络方法,将扩散模型引入水声数据增扩领域,通过构建合成样本集有效提升了现有深度学习模型的泛化能力;提出一种基于扩散模型的混合生成策略,通过预先混合类别标签与数据,扩展数据生成空间,从而进一步降低过拟合风险。在两个公开水声数据集上的实验表明,基于混合扩散模型的数据增扩方法能够合成时频结构逼真的水声数据,其中视觉 Transformer 模型在ShipsEar数据集上准确率最高提升14.6%,有效提高了多种识别模型的分类性能。

关键词: 数据增扩, 水声目标识别, 扩散模型, 深度学习

Abstract:

In addressing the pervasive data scarcity problem in the field of underwater acoustic target recognition, a data augmentation method is proposed based on hybrid diffusion model. Distinguishing itself from conventional frequency spectral transformations or generative adversarial networks methods, the diffusion model is introduced into the field of underwater acoustic data augmentation. By constructing a synthetic sample set, the generalization ability of existing deep learning models is effectively improved. A hybrid generation strategy based on diffusion models is proposed, which expands the data generation space by pre-mixing category labels with data, thereby further reducing the risk of overfitting. Experiments on two public underwater acoustic datasets demonstrate that the data augmentation method based on the hybrid diffusion model can synthesize underwater acoustic data with realistic time-frequency structures. Specifically, the vision Transformer model achieves an accuracy improvement of up to 14.6% on the ShipsEar dataset, effectively enhancing the classification performance of various recognition models.

Key words: data augmentation, underwater acoustic target recognition, diffusion model, deep learning

中图分类号: