Bidirectional Autoregressive Transformer and Fast Fourier Convolution Enhanced Mural Inpainting
Yong CHEN1, 2, Shilong ZHANG1, Wanjun DU1
Author information+
Show less
文章历史+
Received
Published
2024-04-20
2025-04-25
Issue Date
2026-02-12
PDF (9908K)
摘要
针对现有深度学习算法在壁画修复时,存在全局语义一致性约束不足及局部特征提取不充分,导致修复后的壁画易出现边界效应和细节模糊等问题,提出一种双向自回归Transformer与快速傅里叶卷积增强的壁画修复方法.首先,设计基于Transformer结构的全局语义特征修复模块,利用双向自回归机制与掩码语言模型(masked language modeling, MLM),提出改进的多头注意力全局语义壁画修复模块,提高对全局语义特征的修复能力.然后,构建了由门控卷积和残差模块组成的全局语义增强模块,增强全局语义特征一致性约束.最后,设计局部细节修复模块,采用大核注意力机制(large kernel attention, LKA)与快速傅里叶卷积提高细节特征的捕获能力,同时减少局部细节信息的丢失,提升修复壁画局部和整体特征的一致性.通过对敦煌壁画数字化修复实验,结果表明,所提算法修复性能更优,客观评价指标均优于比较算法.
Abstract
Aiming at the lack of global semantic consistency constraints and insufficient acquisition of local features of the current deep learning algorithms in the process of image restoration of broken murals, resulting in the restored murals being prone to boundary effects and blurring of details, this paper proposes a bidirectional autoregressive Transformer with fast Fourier convolutional enhancement of murals restoration method. First, a global semantic feature repair module based on the Transformer structure is designed, and an improved multi-head attention global semantic mural repair module is proposed using the bidirectional autoregressive mechanism with masked language modeling (MLM) to improve the repair capability of global semantic features. Then, a global semantic enhancement module consisting of gated convolution and a residual module is constructed to enhance the global semantic consistency constraint. Finally, the local detail repair module is designed, which adopts large kernel attention (LKA) and fast Fourier convolution (FFC) to improve the ability of capturing detailed features while reducing the loss of local detail information, so as to enhance the consistency of the local and overall features of the repaired murals. The experimental results of the digital restoration of real Dunhuang murals show that the proposed algorithm can effectively restore the structure and texture of the murals, and the subjective visual effect and objective evaluation indexes are better than the comparative algorithms.
在壁画修复过程中,由于普通卷积操作是一种基于局部区域的操作,仅具有局部相关性,其全局特征捕获能力较弱[31],导致壁画修复时易出现边界效应的问题.为了克服上述不足,本文设计了全局语义特征修复模块,利用Transformer结构及双向自回归[32]多头注意力机制建立了壁画像素序列Token之间的依赖关系,并采用掩码语言模型(masked language modeling, MLM)[33]实现对壁画缺失像素的推理,然后通过解码像素序列得到全局语义修复特征图.通过全局语义特征修复模块,提高对全局语义特征的修复能力.
PANY H, LUD M. Digital protection and restoration of Dunhuang mural[J]. Journal of System Simulation,2003,15(3):310-314.(in Chinese)
[3]
WANGH, LIQ Q, JIAS .A global and local feature weighted method for ancient murals inpainting[J].International Journal of Machine Learning and Cybernetics,2020,11(6):1197-1216.
[4]
SCHAEFERK, WEICKERTJ .Diffusion–shock inpainting[M]//Scale Space and Variational Methods in Computer Vision.Cham:Springer International Publishing,2023:588-600.
CHENY, AIY P, GUOH G .Inpainting algorithm for Dunhuang mural based on improved curvature-driven diffusion model[J].Journal of Computer-Aided Design & Computer Graphics,2020,32(5):787-796.(in Chinese)
LIL, GAOR W, MEIS L,et al .Mural image de-noising based on Shannon-Cosine wavelet precise integration method[J].Journal of Zhejiang University (Science Edition),2019,46(3):279-287.(in Chinese)
[9]
BHELES, SHRIRAMWARS, AGARKARP .An efficient texture-structure conserving patch matching algorithm for inpainting mural images[J].Multimedia Tools and Applications,2023,82(30):46741-46762.
WANGH, LIL, LIQ,et al .A global uniform and local continuity repair method for murals inpainting[J].Journal of Hunan University (Natural Sciences),2022,49(6):135-145.(in Chinese)
CHENY, DUW J, ZHAOM X .Improved sparse mural restoration algorithm using joint adaptive learning of multiple dictionaries[J].Journal of Hunan University (Natural Sciences),2023,50(12):1-9.(in Chinese)
[18]
YANGJ, RUHAIYEMN I R, ZHOUC C .A 3 M-hybrid model for the restoration of unique giant murals:a case study on the murals of Yongle Palace[EB/OL]. [2024-04-20]. org/abs/2309.06194v1.
ZHAOL, LINS H, LINZ J,et al .Progressive multilevel feature inpainting algorithm for Chinese ancient paintings[J].Journal of Computer-Aided Design & Computer Graphics,2023,35(7):1040-1051.(in Chinese)
[21]
WADHWAG, DHALLA, MURALAS,et al .Hyperrealistic image inpainting with hypergraphs[C]//2021 IEEE Winter Conference on Applications of Computer Vision (WACV). Waikoloa,HI,USA. IEEE,2021:3911-3920.
[22]
GUOX F, YANGH Y, HUANGD.Image inpainting via conditional texture and structure dual generation[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, QC, Canada. IEEE, 2021: 14114-14123.
[23]
WANGN, LIJ Y, ZHANGL F,et al .MUSICAL:multi-scale image contextual attention learning for inpainting[C]//Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence.Macao,China. 2019: 3748-3754.
[24]
YANGJ, QIZ Q, SHIY .Learning to incorporate structure knowledge for image inpainting[J].Proceedings of the AAAI Conference on Artificial Intelligence,2020,34(7):12605-12612.
ZHAOL, JIB Y, XINGW,et al .Ancient painting inpainting algorithm based on multi-channel encoder and dual attention[J].Journal of Computer Research and Development,2023,60(12):2814-2831.(in Chinese)
[32]
ZENGY, LINZ, LUH C,et al .CR-fill:generative image inpainting with auxiliary contextual reconstruction[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV). Montreal,QC,Canada. IEEE,2021:14144-14153.
[33]
LIW B, LINZ, ZHOUK,et al .MAT:mask-aware transformer for large hole image inpainting[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans,LA,USA. IEEE,2022:10748-10758.
[34]
DENGY, HUIS Q, ZHOUS P,et al .T-former:an efficient transformer for image inpainting[C]//Proceedings of the 30th ACM International Conference on Multimedia. Lisboa,Portugal. ACM, 2022: 6559-6568.
[35]
ZENGY H, FUJ L, CHAOH Y,et al. Aggregated contextual transformations for high-resolution image inpainting[J]. IEEE Transactions on Visualization and Computer Graphics, 2023, 29(7):3266-3280.
WANGZ Y, JIANGS C, SONGQ H,et al. Transformer-based image restoration method for cultural relics[J]. Journal of Computer Research and Development,2024,61(3): 748-761.(in Chinese)
[38]
LIUQ K, TANZ T, CHEND D,et al .Reduce information loss in transformers for pluralistic image inpainting[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans,LA,USA. IEEE,2022:11337-11347.
[39]
PENGZ L, GUOZ H, HUANGW,et al .Conformer:local features coupling global representations for recognition and detection[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence,2023,45(8):9454-9468.
[40]
VASWANIA, SHAZEERN, PARMARN,et al. Attention is all you need[C]//Proceedings of the 31st International Conference on Neural Information Processing Systems.Red Hook, NYCurran Associates Inc,2017:6000-6010.
DEVLINJ, CHANGM W, LEEK T,et al. BERT: pretraining of deep bidirectional transformers for language understanding[C]//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics. Minnesota, MN:Association for Computational Linguistics,2019:4171-4186.
[43]
ZHOUJ H, WEIC, WANGH Y, et al. iBOT:image BERT pre-training with online tokenizer[EB/OL]. [2024-04-20].
[44]
QIUX P, SUNT X, XUY G,et al .Pre-trained models for natural language processing:a survey[J].Science China Technological Sciences, 2020, 63(10): 1872-1897.
[45]
YUJ H, LINZ, YANGJ M, et al. Free-form image inpainting with gated convolution[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). Seoul,Korea (South). IEEE, 2019: 4470-4479.
[46]
LIY H, ZHANGX F, CHEND M. CSRNet:dilated convolutional neural networks for understanding the highly congested scenes[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City,UT, USA. IEEE,2018: 1091-1100.
SUVOROVR, LOGACHEVAE, MASHIKHINA,et al .Resolution-robust large mask inpainting with Fourier convolutions[C]//2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Waikoloa,HI,USA. IEEE,2022:3172-3182.
[49]
JAINJ, ZHOUY Q, YUN, et al. Keys to better image inpainting:structure and texture go hand in hand[C]//2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Waikoloa, HI, USA. IEEE,2023: 208-217.