系统工程与电子技术 ›› 2026, Vol. 48 ›› Issue (10): 3309-3319.doi: 10.12305/j.issn.1001-506X.2026.10.04

• 电子技术 • 上一篇    

多模态大模型驱动的复杂目标毁伤评估方法

李小春(), 葛亚辉(), 任爽   

  1. 空军工程大学信息与导航学院,陕西 西安 710077
  • 收稿日期:2025-07-31 出版日期:2026-10-25 发布日期:2026-09-30
  • 通讯作者: 李小春 E-mail:chunwind@sohu.com;2872214359@qq.com
  • 作者简介:葛亚辉(1986—),男,硕士研究生,主要研究方向为数据智能处理
    任 爽(1998—),女,助教,硕士,主要研究方向为数据智能处理
  • 基金资助:
    学位研究生项目(JY2024C116)资助课题

Multimodal large model-powered method for complex target damage assessment

Xiaochun Li(), Yahui Ge(), Shuang Ren   

  1. Information and Navigation School,Air Force Engineering University,Xi’an 710077,China
  • Received:2025-07-31 Online:2026-10-25 Published:2026-09-30
  • Contact: Xiaochun Li E-mail:chunwind@sohu.com;2872214359@qq.com

摘要:

针对现有毁伤评估技术主要聚焦于提取变化信息、毁伤特征判别难、毁伤推理自动化程度低的问题,借鉴多模态大模型的技术优势,提出一种基于多模态大模型的复杂目标毁伤评估方法。首先,设计基于视觉−语言模型的物理毁伤分析模型,并采用计算机视觉的毁伤语义分析技术设计基于SegGPT的幻觉抑制模型,优化物理部位毁伤分析的准确性。然后,结合毁伤规则知识注入,设计融入毁伤树的毁伤推理模型,并提出一种基于小样本的目标检测识别算法,确保毁伤推理的自动实现。最后,以Markdown形式将毁伤分析和推理的结果以可视化表格形式呈现。实验表明,所提算法可自动实现对复杂目标毁伤部位及毁伤程度的准确判断,物理毁伤评估准确率比Qwen2.5方法提高约29.5%,并能实现从局部到整体的推理能力,同时对不超过4 000×2 492像素的图像,毁伤评估时间在1 min之内,实现了较好的实时性。

关键词: 多模态大模型, 视觉? 语言模型, 毁伤评估, 思维树, 毁伤推理

Abstract:

Aiming at the problems that the existing damage assessment technology mainly focuses on the extraction of change information, the difficulty of damage feature discrimination, and the low degree of automation of damage reasoning, a complex target damage assessment method based on multimodal large scale model is proposed by using the technical advantages of multimodal large model for reference. Firstly, the physical damage analysis model based on visual-language model is designed, and the hallucination suppression model based on SegGPT is designed by using the damage semantic analysis technology of computer vision to optimize the accuracy of physical parts damage analysis. Then, combined with damage rule knowledge injection, a damage reasoning model integrated with damage tree is designed, and a target detection and recognition algorithm based on small samples is proposed to ensure the automatic realization of damage reasoning. Finally, the results of damage analysis and reasoning are presented visual tables in the form of Markdown. Experimental results show that the proposed algorithm can automatically judge the damage position and damage degree of complex targets, which the accuracy of physical damage assessment is improved by approximately 29.5% compared to the Qwen2.5 method, and can achieve the reasoning ability from local to overall. simultaneously, for images of up to 4 000×2 492 pixels, the damage assessment can be completed within 1 min, leading to good real-time performance.

Key words: multimodal large model, vision-language model, damage assessment, tree of thought (ToT), damage reasoning

中图分类号: