基于层级化Token选择Transformer的多模态目标重识别方法

虞智 ,  刘佳文 ,  黄智勇 ,  仲元红

电子科技大学学报 ›› 2026, Vol. 55 ›› Issue (4) : 600 -612.

PDF (5099KB)
电子科技大学学报 ›› 2026, Vol. 55 ›› Issue (4) : 600 -612. DOI: 10.12178/1001-0548.2025230
计算机工程与应用

基于层级化Token选择Transformer的多模态目标重识别方法

作者信息 +

Multimodal object re-identification based on hierarchical token selection Transformer

Author information +
文章历史 +
PDF (5221K)

摘要

多模态目标重识别旨在融合不同模态的互补信息,在视域不重叠的监控场景中实现对同一目标的检索,主要挑战在于缓解跨模态差异。为此,提出一种层级化token选择Transformer(HTSTrans)。HTSTrans由多模态共享特征学习模块(MSFLM)、层级化token选择模块(HTSM)和跨模态对齐约束模块(CACM)构成。MSFLM使用权重共享的视觉Transformer(ViT)学习多模态共享特征,在不同模态间初步建立特征一致性,为后续更精细化的特征学习提供统一的表示空间。HTSM使用层级感知的递进式窗口选择关键token,基于token融合实现对模态特定关键细节特征的提取,再通过多模态特征融合降低跨模态差异。CACM使用模态间传输约束(ITC)和循环一致性约束(CCC)降低同一目标身份在不同模态空间中的特征差异,并提升模态不变性特征的判别性。实验结果表明,HTSTrans可有效提升多模态目标重识别性能。

Abstract

Multimodal object re-identification aims to integrate complementary information from different modals to achieve the retrieval of the same target across non-overlapping surveillances. The primary challenge in this process lies in mitigating the cross-modal discrepancies. To address this issue, a hierarchical token selection transformer (HTSTrans) is proposed. HTSTrans consists of a multi-modal shared feature learning module (MSFLM), a hierarchical token selection module (HTSM), and a cross-modal alignment constraint module (CACM). MSFLM employs a weight-shared vision transformer (ViT) to learn shared multi-modal features, establishing preliminary feature consistency across different modals and providing a unified representation space for subsequent fine-grained feature learning. HTSM utilizes a hierarchically-aware progressive window mechanism to select crucial tokens. Through token fusion, it extracts modal-specific crucial detailed features, followed by multi-modal feature fusion to reduce cross-modal differences. CACM applies an inter-modal transmission constraint (ITC) and a cycle consistency constraint (CCC) to reduce feature discrepancies of the same identity in different modal spaces and enhance the discriminability of modal-invariant features. Experiments demonstrate that HTSTrans effectively improves the performance of multi-modal object re-identification.

关键词

多模态目标重识别 / token选择 / 跨模态对齐 / Transformer

Key words

multi-modal object re-identification / token selection / cross-modal alignment / Transformer

引用本文

引用格式 ▾
虞智,刘佳文,黄智勇,仲元红. 基于层级化Token选择Transformer的多模态目标重识别方法[J]. 电子科技大学学报, 2026, 55(4): 600-612 DOI:10.12178/1001-0548.2025230

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

Qian Yan, Barthelemy J, Karuppiah E, et al. Identifying re-identification challenges: Past, current and future trends[J].SN Computer Science, 2024, 5(7): 937.

[2]

Ye Mang, Chen Shuoyi, Li Chenyue, et al. Transformer for object re-identification: A survey[J].International Journal of Computer Vision, 2025, 133(5): 2410-2440.

[3]

Zhang Shizhou, Luo Wenlong, Cheng De, et al. Prompt-based modality alignment for effective multi-modal object re-identification[J].IEEE Transactions on Image Processing, 2025, 34: 2450-2462.

[4]

Yu Zhi, Huang Zhiyong, Hou Mingyang, et al. Representation selective coupling via token sparsification for multi-spectral object re-identification[J].IEEE Transactions on Circuits and Systems for Video Technology, 2025, 35(4): 3633-3648.

[5]

Wang Yuhao, Lv Yongfeng, Zhang Pingping, et al. IDEA: Inverted text with cooperative deformable aggregation for multi-modal object re-identification[C]//Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025: 29701-29710.

[6]

Guo Jinbo, Zhang Xiaojing, Liu Zhengyi, et al. Generative and attentive fusion for multi-spectral vehicle re-identification[C]//Proceedings of the 7th International Conference on Intelligent Computing and Signal Processing. Xi’an: IEEE, 2022: 1565-1572.

[7]

Zheng Aihua, Zhu Xianpeng, Ma Zhiqi, et al. Cross-directional consistency network with adaptive layer normalization for multi-spectral vehicle re-identification and a high-quality benchmark[J].Information Fusion, 2023, 100: 101901.

[8]

Yan Tianying, Ma Huixin, Wang Changhai, et al. Generalizable multi-spectral vehicle re-identification via decoupled subspaces[C]//Proceedings of the Second International Conference on Applied Intelligence. Zhengzhou: Springer, 2024: 13-24.

[9]

Wang Zi, Li Chenglong, Li Pengyu, et al. Prototype-based diversity and integrity learning for all-day multi-modal person re-identification[J].IEEE Transactions on Information Forensics and Security, 2025, 20: 11385-11400.

[10]

Pan Wenjie, Wu Hanxiao, Zhu Jianqing, et al. H-ViT: Hybrid vision transformer for multi-modal vehicle re-identification[C]//Proceedings of the 2nd CAAI International Conference on Artificial Intelligence. Beijing: Springer, 2022: 255-267.

[11]

Pan Wenjie, Huang Linhan, Liang Jianbao, et al. Progressively hybrid transformer for multi-modal vehicle re-identification[J].Sensors, 2023, 23(9): 4206.

[12]

Wang Yuhao, Liu Xuehu, Zhang Pingping, et al. TOP-ReID: Multi-spectral object re-identification with token permutation[C]//Proceedings of the 38th AAAI Conference on Artificial Intelligence. Vancouver: AAAI, 2024: 5758-5766.

[13]

Zheng Aihua, Ma Zhiqi, Sun Yongqi, et al. Flare-aware cross-modal enhancement network for multi-spectral vehicle Re-identification[J].Information Fusion, 2025, 116: 102800.

[14]

Zheng Aihua, Wang Zi, Chen Zihan, et al. Robust multi-modality person re-identification[C]//Proceedings of 35th AAAI Conference on Artificial Intelligence. Vancouver: AAAI, 2021: 3529-3537.

[15]

Li Hongchao, Li Chenglong, Zhu Xianpeng, et al. Multi-spectral vehicle re-identification: A challenge[C]//Proceedings of 34th AAAI Conference on Artificial Intelligence. New York: AAAI, 2020: 11345-11353.

[16]

Wang Zi, Li Chenglong, Zheng Aihua, et al. Interact, embed, and EnlargE: Boosting modality-specific representations for multi-modal person re-identification[C]//Proceedings of the 36th AAAI Conference on Artificial Intelligence. Vancouver: AAAI, 2022: 2633-2641.

[17]

Zheng Aihua, He Ziling, Wang Zi, et al. Dynamic enhancement network for partial multi-modality person re-identification[PP/OL].V1. arXiv (2023-05-25)[2025-05-25].https://arxiv.org/abs/2305.15762.

[18]

Wu Di, Liu Zhihui, Chen Zihan, et al. LRMM: Low rank multi-scale multi-modal fusion for person re-identification based on RGB-NI-TI[J].Expert Systems with Applications, 2025, 263: 125716.

[19]

Yang Xi, Dong Wenjiao, Cheng De, et al. TIENet: A tri-interaction enhancement network for multimodal person reidentification[J].IEEE Transactions on Neural Networks and Learning Systems, 2025, 36(6): 9852-9863.

[20]

Crawford J, Yin Haoli, Mcdermott L, et al. UniCat: Crafting a stronger fusion baseline for multimodal re-identification[PP/OL]. V1.arXiv (2023-10-28)[2025-06-28].https://arxiv.org/abs/2310.18812.

[21]

Zhang Pingping, Wang Yuhao, Liu Yang, et al. Magic tokens: Select diverse tokens for multi-modal object re-identification[C]//Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 17117-17126.

[22]

Yu Zhi, Huang Zhiyong, Hou Mingyang, et al. WTSF-ReID: Depth-driven window-oriented token selection and fusion for multi-modality vehicle re-identification with knowledge consistency constraint[J].Expert Systems with Applications, 2025, 274: 126921.

基金资助

国家自然科学基金(62372074)

重庆市自然科学基金(CSTB2023NSCQ-MSX0274)

中央高校基本科研基金(2023CDJKYJH050)

AI Summary AI Mindmap
PDF (5099KB)

61

访问

0

被引

详细

导航
相关文章

AI思维导图

/