School of Information Science and Technology, Xizang University, Lhasa 850000, China
Show less
文章历史+
Received
Published
2025-09-23
2026-03-15
Issue Date
2026-07-06
PDF (1768K)
摘要
文章提出一种融合轻量化(Ghost Transformer)结构与动态可变形卷积机制结合的西藏壁画破损检测模型,旨在解决传统方法在壁画裂缝、脱落等破损目标检测中精度低、实时性差的问题,提升壁画破损检测的智能化与高效化水平,为文物数字化保护提供技术支撑。该模型以Ghost Transformer 作为主干网络,搭配BiFPN特征融合结构,并集成Coordinate Attention(CA)与 Sim AM 无参注意力机制,同时引入TOOD动态标签分配策略,通过多技术融合增强模型对小目标与复杂背景的适应能力,实现对西藏壁画破损目标的精准识别与检测。实验结果表明,在检测精度维持稳定的情况下,网络深度缩短31.37%,参数规模下降33.98%,浮点计算量削减50.32%;模型体积压缩至原始体积的35.57%,整体压缩率达64.43%。改进后的GT-D-CA-S-B-T YOLOv8兼顾精度检测与轻量化水平,能有效助力西藏壁画数字化保护,显著提升传统保护修复工作中破损检测的效率和准确性。
Abstract
Aiming to resolve the problems of low accuracy and poor real-time performance in traditional methods for detecting damage, such as cracks and peeling, in ancient murals, this paper presents a model specifically designed for detecting damage on wall murals in Xizang. The proposed model integrates a lightweight Ghost Transformer architecture with a dynamic deformable convolution mechanism. To improve the intelligence and efficiency of mural damage detection and provide technical support for the digital preservation of cultural heritage, the model employs the Ghost Transformer as its backbone network and incorporates a BiFPN feature fusion architecture. Additionally, it incorporates both the Coordinate Attention (CA) and Sim AM parameter-free attention mechanisms, along with the TOOD dynamic label assignment strategy. Through the integration of these multiple technologies, the model enhances its adaptability to small targets and complex backgrounds, enabling more precise identification and detection of damage on wall murals in Xizang. Experimental results show that the improved model offers significant lightweight advantages. While maintaining stable detection accuracy, it reduces network depth by 31.37%, decreases parameter size by 33.98%, and cuts floating-point computations by 50.32%. The model size is reduced to 35.57% of its original volume, resulting in an overall compression rate of 64.43%. The improved GT-D-CA-S-B-T YOLOv8 effectively balances detection accuracy with lightweight performance, providing robust technical support for the digital preservation of Xizang’s wall murals and significantly improving the efficiency and accuracy of damage detection in traditional conservation and restoration work.
壁画是集艺术、历史与科学研究价值于一体的重要文化遗产,承载着区域历史文化、宗教信仰与民族记忆[1]。壁画广泛分布于我国西部及边疆地区的寺庙、洞窟等文物遗址中。西藏壁画作为藏族优秀传统文化传承的重要载体,不仅凝聚着古代工匠的艺术匠心与工艺智慧,更以图像史的形式记录了区域历史的发展脉络。然而,受自然风化、地震扰动、微生物侵蚀以及人为破坏等多种因素影响,壁画本体普遍出现裂缝、起鼓、剥落等破损现象。传统文物的检测与修复工作长期依赖人工经验判断,不仅效率低下,更受限于人力与时间成本的双重限制,难以支撑大规模文物普查与系统性保护工程的推进。鉴于此,构建一套高效、精准且智能化的壁画破损检测技术体系已成为亟待解决的问题,对于推动我国文化遗产保护工作向数字化、智能化方向发展具有重要意义。在此背景下,基于深度学习的图像检测技术为壁画破损的自动识别提供了可行路径。近年来以YOLO(You Only Look Once)系列为代表的单阶段检测网络在目标检测领域的效率与精度上取得显著进展,为构建实时、高性能壁画破损检测模型提供了技术基础。然而,传统YOLO结构在检测小目标(如壁画的细裂缝、局部脱落)及高纹理背景(如复杂壁画图案)目标时存在漏检和误检的问题。此外,标准卷积结构对不规则边缘和细微破损的表征能力有限。因此,如何将轻量化网络结构与多尺度、变形感知能力相结合,以提升模型对壁画破损目标的检测性能,已成为当前研究中亟待解决的关键问题。当前,国际文化遗产领域的数字化工作主要聚焦于几何结构与纹理信息的记录。如,欧洲的CHRESP项目与意大利的DigiArt项目,均侧重于文物的三维建模与重建[2],而对于壁画本体的破损检测方面仍以人工分析为主。近年来,卷积神经网络(CNN)等深度神经网络广泛用于工业检测与医学图像分割任务,并逐步扩展至文化遗产领域。例如,FasterR-CNN被用于裂缝定位[3],YOLOv5被尝试应用于西藏壁画脱落检测[4],但其对细节特征的建模仍然存在一定的局限性。
在特征提取阶段,模型将坐标注意力(CA)机制与无参数注意力(Sim AM)机制进行串联融合,构建了双注意力协同增强框架。其中,CA通过方向感知分解与位置编码策略,将二维空间注意力解耦为水平与垂直方向的特征交互,能够精准建模壁画表面裂缝、剥落等破损区域的空间位置关联性,有效抑制背景冗余信息干扰;而Sim AM 则基于能量函数优化理论,通过局部能量响应值,在不引入额外参数的前提下,增强模型对低对比度、小目标破损区域的特征响应灵敏度,从而有效缓解传统注意力机制在轻量化场景下对细微破损特征捕捉不足的问题。
HouY, KenderdineS, PiccaD, et al. Digitizing intangible cultural heritage embodied: State of the art[J]. Journal on Computing and Cultural Heritage (JOCCH), 2022, 15(3): 1-20.
[3]
ShaoqingR, KaimingH, RossG, et al.Faster R-CNN: Towards real-time object detection with region proposal networks[J].IEEE transactions on pattern analysis and machine intelligence,2017,39(6):1137-1149.
TanM, PangR, LeQ V. EfficientDet: Scalable and efficient object detection[C].Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition,2020:10778-10787.
[9]
DaiJ, QiH, XiongY, et al. Deformable convolutional networks[C]. Proceedings of the IEEE International Conference on Computer Vision, 2017:764-773.
[10]
HouQ, ZhouD, FengJ. Coordinate attention for efficient mobile network design[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021:13713-13722.
[11]
LiuQ, HuangW, DuanX, et al. DSW-YOLOv8n: A new underwater target detection algorithm based on improved YOLOv8n[J]. Electronics, 2023,12(18): 3892.
[12]
FengC, ZhongY, GaoY, et al. Tood: Task-aligned one-stage object detection[C]. 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE Computer Society, 2021:3490-3499.
[13]
LiL, WangZ, Gbh-yolovT Z. Ghost convolution with bottleneckcsp and tiny target prediction head incorporating yolov5 for pv panel defect detection.[J].Electronics,2023,12(3), 561-576
[14]
ZhangJ, LiX, LiJ, et al. Rethinking mobile block for efficient attention-based models[C]. 2023 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE Computer Society, 2023: 1389-1400.
[15]
HanK, WangY, TianQ, et al. Ghostnet: More features from cheap operations[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020: 1580-1589.
[16]
TanH, LiuX, YinB, et al. MHSA-Net: Multihead self-attention network for occluded person re-identification[J]. IEEE Transactions on Neural Networks and Learning Systems, 2022, 34(11): 8210-8224.
[17]
ZhuX, HuH, LinS, et al. Deformable convnets v2: More deformable, better results[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019: 9308-9316.
向寅鸿,周恺卿, SARKHEYLI-H?GELEArezoo,等.基于可逆和动态分解机制的层次化FPN并行故障诊断(Parallel fault diagnosis using hierarchical fuzzy Petri net byreversible and dynamic -decomposition mechanism)[J].Frontiers of Information Technology & Electronic Engineering,2025,26(1):93-109.
BaojunZ, BoyaZ, LinboT, et al. Multi-scale object detection by top-down and bottom-up feature pyramid network[J].Journal of Systems Engineering and Electronics,2019,30(1):1-12.
[22]
李军,孟佳兵,李攀.利用优化的YOLOv7模型自动检测地震反演低频模型中的“牛眼”效应(Detecting the Bull's-Eye Effect in Seismic Inversion Low-Frequency Models Using the Optimized YOLOv7 Model)[J].Applied Geophysics,2024,21(4):766-776+880-881.