A text-guided face image inpainting method is proposed to address the problems of structural distortion, texture blurring, and uncontrollability in current face inpainting methods. The method reconstructs the missing regions in an image by fusing image features and corresponding text features. In the network training, a visual-textual modal fusion module is designed for associating image and textual features so that the reconstruction of missing regions of the face is not only based on the visual semantics visible in the image, but also guided by textual semantics with rich text. An attention-aware layer is added between the encoded and decoded features to improve the consistency of the appearance of the visible and generated regions. Experimental results on the CelebA-HQ face dataset show that the method in this paper is able to obtain restoration results that are more natural and consistent with the textual semantics in terms of texture and structure, and its visual effect and evaluation metrics are better than those of the comparison algorithms.
ZhouDa-ke, ZhangChao, YangXin.Self-supervised 3D face reconstruction based on multi-scale feature fusion and dual attention mechanism[J]. Journal of Jilin University (Engineering and Technology Edition), 2022, 52(10): 2428-2437.
WangXiao-yu, HuXin-hao, HanChang-lin. Face pencil drawing algorithms based on generative adversarial network[J]. Journal of Jilin University (Engineering and Technology Edition), 2021, 51(1): 285-292.
[5]
PathakD, KrahenbuhlP, DonahueJ, et al. Context encoders: feature learning by inpainting[C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, 2016: 2536-2544.
[6]
IizukaS, SimoS E, IshikawaH. Globally and locally consistent image completion[J]. ACM Transactions on Graphics (ToG), 2017, 36(4): 1-14.
[7]
YanZ, LiX, LiM, et al. Shift-net: image inpainting via deep feature rearrangement[C]∥Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 2018: 1-17.
[8]
LiuH, WanZ, HuangW, et al. Pd-gan: probabilistic diverse gan for image inpainting[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 9371-9381.
[9]
WanZ, ZhangJ, ChenD, et al. High-fidelity pluralistic image completion with transformers[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision, Nashville, USA, 2021: 4692-4701.
[10]
LiW, LinZ, ZhouK, et al. Mat: mask-aware transformer for large hole image inpainting[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 10758-10768.
[11]
HuangW, DengY, HuiS, et al. Sparse self-attention transformer for image inpainting[J]. Pattern Recognition, 2024, 145: 109897.
[12]
LianJ, ZhangJ, LiuJ, et al. Guiding image inpainting via structure and texture features with dual encoder[J]. The Visual Computer, 2024, 40: 4303-4317.
[13]
DevlinJ, ChangM W, LeeK, et al. Bert: pre-training of deep bidirectional transformers for language understanding[J/OL]. [2023-12-16]. arXiv preprint arXiv:
[14]
JohnsonJ, AlahiA, FeiF L. Perceptual losses for real-time style transfer and super-resolution[C]∥14th European Conference, Amsterdam, The Netherlands, 2016: 694-711.
[15]
RussakovskyO, DengJ, SuH, et al. Imagenet large scale visual recognition challenge[J]. International Journal of Computer Vision, 2015, 115: 211-252.
[16]
SimonyanK, ZissermanA. Very deep convolutional networks for large-scale image recognition[J/OL].[2023-12-17]. arXiv preprint arXiv:
[17]
MaoX, LiQ, XieH, et al. Least squares generative adversarial networks[C]∥Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 2017: 2794-2802.
[18]
GatysL A, EckerA S, BethgeM. Image style transfer using convolutional neural networks[C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, 2016: 2414-2423.
[19]
LiuG, RedaF A, ShihK J, et al. Image inpainting for irregular holes using partial convolutions[C]∥Proceedings of the European Conference on Computer Vision, Munich, Germany, 2018: 85-100.
[20]
LiJ, WangN, ZhangL, et al. Recurrent feature reasoning for image inpainting[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 7760-7768.
[21]
LugmayrA, DanelljanM, RomeroA, et al. Repaint: inpainting using denoising diffusion probabilistic models[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 11461-11471.
[22]
ChenL, YuanC, QinX, et al. Contrastive structure and texture fusion for image inpainting[J]. Neurocomputing, 2023, 536: 1-12.
[23]
LiA, ZhaoL, ZuoZ, et al. MIGT: multi-modal image inpainting guided with text[J]. Neurocomputing, 2023, 520: 376-385.
[24]
ZhangL, ChenQ, HuB, et al. Text-guided neural image inpainting[C]∥Proceedings of the 28th ACM International Conference on Multimedia, New York, USA, 2020: 1302-1310.