文图跨模态知识蒸馏的行人检索算法
Research on person re trieval algorithm based on text-to-image cross-modal knowledge distillation
当前文图跨模态行人检索算法过于依赖视觉语言预训练大模型来提升精度,导致模型参数量庞大、算力要求高等问题,难以满足边缘部署等实际应用需求。基于此,该文从视觉语言预训练大模型的轻量化角度入手,引入多模态联合蒸馏与互补式监督策略,基于华为昇腾平台,提出了一种基于三阶段渐进式知识蒸馏的文图跨模态行人检索算法。该方法区别于传统两阶段仅对各个模态进行独立蒸馏,而是通过模态内−跨模态的蒸馏路径,依次在图像、文本模态内实现特征对齐,最终在共享隐空间中进行跨模态语义关联的协同蒸馏。实验表明,学生模型参数量仅为教师模型的14.77%,在CUHK-PEDES、ICFG-PEDES和RSTPReid数据集上与现有轻量化方法相比,其mAP指标均得到有效提升,同时在边缘设备上处理单个文本−图像对的耗时约23 ms,实现了实时推理性能。该研究证实了基于文图跨模态知识蒸馏的行人检索算法的有效性,为人员排查、智能安防等场景国产化落地提供了有效路径。
Current text-to-image cross-modal person retrieval algorithms overly rely on vision-language pre-trained large models to improve accuracy. This leads to issues such as large model parameter sizes and high computational requirements, making it difficult to meet the practical application requirements such as edge deployment. In light of this, this paper focuses on the lightweight aspect of vision-language pre-trained large models. It introduces a multi-modal joint distillation and complementary supervision strategy. Based on Huawei Ascend platform, the paper proposes a text-image cross-modal pedestrian retrieval algorithm based on a three-stage progressive knowledge distillation approach. Unlike the traditional two-stage approach that only performs independent distillation for each modality, this method first conducts distillation within the image and text modalities separately, and finally performs collaborative distillation for cross-modal semantic association in the shared latent space. Experimental results show that the student model has only 14.77% of the parameters of the teacher model. Compared with existing lightweight methods, it achieves effective improvements in mAP on the CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets. It also realizes real-time inference performance with an inference time of approximately 23 ms for processing a single text-image pair on edge devices. This study confirms the effectiveness of the three-stage progressive knowledge distillation method and provides an effective pathway for the domestic implementation in scenarios such as personnel screening and intelligent security.
| [1] |
|
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
|
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
石瑞鑫, 智敏, 殷雁君 . 多模态行人重识别研究综述[J].计算机应用研究, 2025, 42(7): 1921-1929. |
| [16] |
|
| [17] |
|
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
|
| [22] |
|
| [23] |
|
| [24] |
邵仁荣, 刘宇昂, 张伟, |
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
|
| [29] |
|
| [30] |
|
| [31] |
|
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
|
| [36] |
|
| [37] |
|
| [38] |
|
| [39] |
|
| [40] |
|
| [41] |
|
国家自然科学基金(62372082)
中央高校基本科研业务费(ZYGX2024Z017)
深圳市自然科学基金(JCYJ20240813114206010)
/
| 〈 |
|
〉 |