Systems Engineering and Electronics ›› 2026, Vol. 48 ›› Issue (5): 1481-1491.doi: 10.12305/j.issn.1001-506X.2026.05.04

• Electronic Technology • Previous Articles     Next Articles

Lightweight object detection method based on visible and infrared feature fusion

Jie ZHANG1,*(), Tianqing CHANG2, Xiaowei WANG1, Wenlong HAO1, Xin TANG1   

  1. 1. Army Aviation Institute,Beijing 101123,China
    2. Army Arms University of PLA,Beijing 100072,China
  • Received:2025-03-19 Accepted:2025-07-22 Online:2026-05-27 Published:2026-05-27
  • Contact: Jie ZHANG E-mail:zjwhy_8@163.com

Abstract:

To address the problems of low efficiency and high computational complexity of dual-stream object detection models, a lightweight object detection method based on visible and infrared feature fusion is proposed. Firstly, you only look once (YOLO) v8 is expanded into a dual-stream object detection model, the dual-stream backbone network is optimized using group convolution, and the two independent backbone networks are merged into one backbone network, which realizes the synchronous extraction of two modal features, greatly improving the model operation efficiency. Secondly, the faster cross-stage partial bottleneck with two convolution with cross-modal feature interaction (C2f-CMFI) module and the spatial pyramid pooling fast with cross-modal feature fusion (SPPF-CMFF) module are designed, while reducing the complexity of the model, fusion and interaction of the two modal features during the feature extraction process are realized. Finally, the experimental results on the public visible-infrared dataset show that compared with the traditional dual-stream object detection models, the parameter amount and computational complexity of the proposed method are reduced by 19.5% and 17.7% respectively, and the mean average precision 50:95 is improved by 1.9%. On a NVIDIA RTX 2080Ti graphics processing unit, the inference speed is 140 frames per second, which proved the effectiveness of the proposed method.

Key words: visible-infrared image, you only look once (YOLO), lightweighting, object detection, dual-stream structure

CLC Number: 

[an error occurred while processing this directive]