基于模态度量的多模态海上船舶识别方法

徐政伟 ,  黄沛基 ,  何建樑 ,  王世豪

河南师范大学学报(自然科学版) ›› 2026, Vol. 54 ›› Issue (5) : 105 -111.

PDF (4331KB)
河南师范大学学报(自然科学版) ›› 2026, Vol. 54 ›› Issue (5) : 105 -111. DOI: 10.16366/j.cnki.1000-2367.2025.06.25.0001
数学与计算机科学

基于模态度量的多模态海上船舶识别方法

作者信息 +

Multimodal maritime ship recognition method based on modality metric

Author information +
文章历史 +
PDF (4434K)

摘要

现有的海上多模态识别方法在处理复杂非线性特征方面存在欠缺之处,无法充分利用不同模态的信息,而且融合过程采用的是固定融合策略、融合权重.基于此提出了一种基于模态度量的海上船舶识别网络,这个模型引进了特征过滤模块,能够抑制无关信息,减轻模态间的分布差异,增强多模态特征提取的鲁棒性.本文还设计了模态度量评估机制,依据各模态特征的贡献来动态分配融合权重,从而最大程度地减少低价值模态对最终融合性能的影响.此机制能够根据当前场景、环境的变化,自适应地调整融合权重,提高融合效果.本文提出的方法在公开的VAIS海上红外可见光多模态数据集上进行了基准测试,在准确性和计算效率方面都优于当前最先进的方法.

Abstract

To address the shortcomings that current maritime multimodal recognition methods have in processing complex nonlinear features, effectively utilizing information from various modalities, and using fixed fusion strategies and weights, this paper puts forward a Modality Measurement-based Maritime Ship Recognition Network (MM-NET). The model has inside a Feature Filtering Module (FFM) which is used for effective suppression of useless information, reduction of distribution differences between different modalities, and improvement of the stability of multimodal feature picking work. In addition, this article devises a Modality Measurement (MM) mechanism which dynamically gives out fusion weights according to the contribution of every modal feature, therefore reducing to the smallest extent the influence that low-value modalities have on the final fusion performance. This mechanism can promote fusion effect degree on current scenarios and enviroment changes, and on self-adaption fusion weight changes. The method we put forward has carried on the benchmark test on the public VAIS maritime infrared-visible multimodal dataset, thus it has shown that the accuracy and calculation efficiency outperform than the existing advanced methods.

关键词

多模态融合 / 海上船舶识别 / 特征过滤 / 模态度量

Key words

multimodal fusion / maritime ship recognition / feature filtering / modality metric

引用本文

引用格式 ▾
徐政伟,黄沛基,何建樑,王世豪. 基于模态度量的多模态海上船舶识别方法[J]. 河南师范大学学报(自然科学版), 2026, 54(5): 105-111 DOI:10.16366/j.cnki.1000-2367.2025.06.25.0001

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

郑海君, 葛斌, 夏晨星, 等 . 多特征聚合的红外-可见光行人重识别[J]. 光电工程, 2023, 50(7): 230136.

[2]

Zheng H J, Ge B, Xia C X, et al. Infrared-visible person re-identification based on multi feature aggregation[J]. Opto-Electronic Engineering, 2023, 50(7): 230136.

[3]

Gao X Y, Shi Y B, Zhu Q, et al. Infrared and visible image fusion with deep neural network in enhanced flight vision system[J]. Remote Sensing, 2022, 14(12): 2789.

[4]

Dong S K, Feng J F, Fang D X . A novel multiscale contrastive learning network for fine-grained ocean ship classification[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2024, 17: 9989-10005.

[5]

Xiao Q, Liu B, Li Z Y, et al. Progressive data augmentation method for remote sensing ship image classification based on imaging simulation system and neural style transfer[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021, 14: 9176-9186.

[6]

He J L, Chang W L, Wang F P, et al. Multi-scale dense networks for ship classification using dual-polarization SAR images[C]// 2023 IEEE Radar Conference (RadarConf23). San Antonio: IEEE, 2023: 1-6.

[7]

杨颖, 杨艳秋, 余本功. 考虑视图可信度的用户多模态意图识别方法[J]. 电子与信息学报, 2025, 47(6): 1966-1975.

[8]

Yang Y, Yang Y Q, Yu B G. Multimodal intent recognition method with view reliability[J]. Journal of Electronics & Information Technology, 2025, 47(6): 1966-1975.

[9]

Zhang M M, Choi J, Daniilidis K, et al. VAIS: a dataset for recognizing maritime imagery in the visible and infrared spectrums[C]// 2015 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Boston: IEEE, 2015: 10-16.

[10]

Liu J Y, Fan X, Huang Z B, et al. Target-awared dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection[C]// 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans: IEEE, 2022: 5792-5801.

[11]

Tlig M, Bouchouicha M, Sayadi M, et al. Infrared-visible images' fusion techniques for forest fire monitoring[C]// 2022 6th International Conference on Advanced Technologies for Signal and Image Processing (ATSIP). Sfax: IEEE, 2022: 1-6.

[12]

Tang L F, Zhang H, Xu H, et al. Rethinking the necessity of image fusion in high-level vision tasks: a practical infrared and visible image fusion network based on progressive semantic injection and scene fidelity[J]. Information Fusion, 2023, 99: 101870.

[13]

Xie Z H, Zong S, Li Q, et al. Interactive residual coordinate attention and contrastive learning for infrared and visible image fusion in triple frequency bands[J]. Scientific Reports, 2024, 14: 90.

[14]

Zhao Z X, Bai H W, Zhang J S, et al. CDDFuse: correlation-driven dual-branch feature decomposition for multi-modality image fusion[C]// 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver: IEEE, 2023: 5906-5916.

[15]

Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: transformers for image recognition at scale[EB/OL]. [2025-06-10]. https://arxiv.org/abs/2010.11929.

[16]

He K M, Zhang X Y, Ren S Q, et al. Deep residual learning for image recognition[C]// 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas: IEEE, 2016: 770-778.

[17]

Yu W H, Wang X C . MambaOut: do we really need mamba for vision[C]// 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville: IEEE, 2025: 4484-4496.

[18]

Liu Z, Mao H Z, Wu C Y, et al. A ConvNet for the 2020s[C]// 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans: IEEE, 2022: 11966-11976.

[19]

Peng X K, Wei Y K, Deng A D, et al. Balanced multimodal learning via on-the-fly gradient modulation[C]// 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans: IEEE, 2022: 8228-8237.

[20]

He X Y, Wang Y, Zhao S, et al. Co-attention fusion network for multimodal skin cancer diagnosis[J]. Pattern Recognition, 2023, 133: 108990.

[21]

Zhang Q, Wu H, Zhang C, et al. Provable dynamic fusion for low-quality multimodal data[C]// International conference on machine learning. [S. l.]: PMLR, 2023: 41753-41769.

[22]

Chitta K, Prakash A, Jaeger B, et al. TransFuser: imitation with transformer-based sensor fusion for autonomous driving[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(11): 12878-12895.

[23]

Vaezi Joze H R, Shaban A, Iuzzolino M L, et al. MMTM: multimodal transfer module for CNN fusion[C]// 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). [S. l.]: IEEE, 2020: 13286-13296.

基金资助

国家自然科学基金(62271499)

河南省科技攻关计划项目(252102211049)

AI Summary AI Mindmap
PDF (4331KB)

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/

〈 〉