基于注意力机制的草图简化

黄昭辉 ,  姚争为

杭州师范大学学报(自然科学版) ›› 2026, Vol. 25 ›› Issue (3) : 320 -328.

PDF (2111KB)
杭州师范大学学报(自然科学版) ›› 2026, Vol. 25 ›› Issue (3) : 320 -328. DOI: 10.19926/j.cnki.issn.1674-232X.2024.09.011
数学与信息科学

基于注意力机制的草图简化

作者信息 +

Attention mechanism-based sketch simplification

Author information +
文章历史 +
PDF (2160K)

摘要

基于漫画创作和产品设计等领域对简化粗糙手绘草图以生成干净线条图的需求,针对现有草图简化方法生成的线条图风格单一、难以保留原有绘画风格的问题,提出了一种草图简化方法,旨在基于参考草图的风格对粗糙草图进行简化,使生成草图在保留原始视觉内容的同时,呈现出与参考草图近似的线条风格.该方法采用数据增强技术扩展数据集,以应对可用于训练的粗糙草图及其对应的简化图像数量有限的问题;在网络中引入卷积块注意力模块(convolutional block attention module,CBAM)和空间与压缩注意力(shifted sparse attention,S2Attention)模块,用于检测粗糙草图和参考草图的边缘特征,并学习二者之间的视觉对应关系;为确保生成图像的质量及对参考线条风格的模仿效果,使用重建损失、线条损失、感知损失和对抗性损失函数对模型进行优化.定性和定量实验结果表明,该方法在模仿参考草图的线条风格方面表现突出,且简化效果优于现有基线方法.

Abstract

Motivated by the need for simplifying rough hand-drawn sketches into clean line art in domains such as comic creation and product design, this paper proposes a novel sketch simplification method to overcome the limitations of existing approaches, which often produce line art of a single style and fail to preserve the original drawing style. This method simplifies a rough sketch guided by the style of a reference sketch, enabling the output to retain the original visual content while adopting a line style akin to that of the reference. To address the scarcity of paired training data consisting of rough sketches and their corresponding simplified images, data augmentation techniques are employed to enlarge the dataset. The network incorporates a convolutional block attention module (CBAM) and a shifted sparse attention (S2Attention) module to extract edge features from both the rough and reference sketches and to learn the visual mapping between them. To guarantee the quality of the generated images and faithful reproduction of the reference line style, the model is optimized using a combination of reconstruction loss, line loss, perceptual loss, and adversarial loss. Both qualitative and quantitative experimental results demonstrate that the proposed method excels at mimicking the line style of reference sketches and outperforms existing baseline methods in terms of simplification quality.

关键词

草图简化 / 图像到图像的转换 / 风格迁移 / 注意力机制

Key words

sketch simplification / image-to-image translation / style transfer / attention mechanism

引用本文

引用格式 ▾
黄昭辉,姚争为. 基于注意力机制的草图简化[J]. 杭州师范大学学报(自然科学版), 2026, 25(3): 320-328 DOI:10.19926/j.cnki.issn.1674-232X.2024.09.011

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

LIU X T, WONG T T, HENG P A. Closure-aware sketch simplification[J]. ACM Transactions on Graphics, 2015, 34(6): 1-10.

[2]

PENG J Y, WANG J X, WANG J, et al. A relic sketch extraction framework based on detail-aware hierarchical deep network[J]. Signal Processing, 2021, 183: 108008.

[3]

LI C Z, LIU X T, WONG T T. Deep extraction of Manga structural lines[J]. ACM Transactions on Graphics, 2017, 36(4): 1-12.

[4]

LIU X T, MAO X Y, YANG X, et al. Stereoscopizing cel animations[J]. ACM Transactions on Graphics, 2013, 32(6): 1-10.

[5]

SASAKI K, IIZUKA S, SIMO-SERRA E, et al. Joint gap detection and inpainting of line drawings[C]// 2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu: IEEE, 2017: 5768-5776.

[6]

BAE S H, BALAKRISHNAN R, SINGH K. I Love Sketch: as-natural-as-possible sketching system for creating 3D curve models[C]// Proceedings of the 21st Annual ACM Symposium on User Interface Software and Technology. New York: ACM, 2008: 151-160.

[7]

ARVO J , NOVINS K. Fluid sketches: continuous recognition and morphing of simple hand-drawn shapes[C]// Proceedings of the 13th Annual ACM Symposium on User Interface Software and Technology. San Diego: ACM, 2000: 73-80.

[8]

PREIM B, STROTHOTTE T. Tuning rendered line-drawings[J]. Journal of WSCG, 1995, 3(1/2): 228-238.

[9]

WILSON B, MA K L. Rendering complexity in computer-generated pen-and-ink illustrations[C]// Proceedings of the 3rd International Symposium on Non-Photorealistic Animation and Rendering. Annecy: ACM, 2004: 129-137.

[10]

COLE F, DECARLO D, FINKELSTEIN A, et al. Directing gaze in 3D models with stylized focus[C]// Proceedings of the 17th Eurographics Conference on Rendering Techniques. Nicosia: ACM, 2006: 377-387.

[11]

GRABLI S, DURAND F, SILLION F X. Density measure for line-drawing simplification[C]// 12th Pacific Conference on Computer Graphics and Applications. Seoul: ACM, 2004: 309-318.

[12]

QI Y G, SONG Y Z, XIANG T, et al. Making better use of edges via perceptual grouping[C]// 2015 IEEE Conference on Computer Vision and Pattern Recognition. Boston: IEEE, 2015: 1856-1865.

[13]

GRIMM C, JOSHI P. Just draw it: a 3D sketching system[C]// International Symposium on Sketch-Based Interfaces and Modeling. Goslar: The Eurographics Association, 2012: 121-130.

[14]

FIŠER J, ASENTE P, SYKORA D. ShipShape: a drawing beautification assistant[C]// International Symposium on Sketch-Based Interfaces and Modeling. Goslar: The Eurographics Association, 2015: 49-57.

[15]

ORBAY G, KARA L B. Beautification of design sketches using trainable stroke clustering and curve fitting[J]. IEEE Transactions on Visualization and Computer Graphics, 2011, 17(5): 694-708.

[16]

FAVREAU J D, LAFARGE F, BOUSSEAU A. Fidelity vs. simplicity: a global approach to line drawing vectorization[J]. ACM Transactions on Graphics, 2016, 35(4): 120.

[17]

SIMO-SERRA E, IIZUKA S, SASAKI K, et al. Learning to simplify: fully convolutional networks for rough sketch cleanup[J]. ACM Transactions on Graphics, 2016, 35(4): 121.

[18]

SIMO-SERRA E, IIZUKA S, ISHIKAWA H. Mastering sketching: adversarial augmentation for structured prediction[J]. ACM Transactions on Graphics, 2018, 37(1): 1-13.

[19]

XU X M, XIE M S, MIAO P Q, et al. Perceptual-aware sketch simplification based on integrated VGG layers[J]. IEEE Transactions on Visualization and Computer Graphics, 2021, 27(1): 178-189.

[20]

SHAHAM T R, GHARBI M, ZHANG R, et al. Spatially-adaptive pixelwise networks for fast image translation[C]// 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021: 14877-14886.

[21]

ZHOU X R, ZHANG B, ZHANG T, et al. CoCosNet v2: full-resolution correspondence learning for image translation[C]// 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021: 11460-11470.

[22]

PARK T, EFROS A A, ZHANG R, et al. Contrastive learning for unpaired image-to-image translation[C]// 16th European Conference on Computer Vision. Cham: Springer, 2020: 319-345.

[23]

NIZAN O, TAL A. Breaking the cycle-colleagues are all you need[C]// 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 7857-7866.

[24]

CHOI Y, UH Y, YOO J, et al. StarGAN v2: diverse image synthesis for multiple domains[C]// 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020: 8185-8194.

[25]

PARK T, ZHU J Y, WANG O, et al. Swapping autoencoder for deep image manipulation[J]. Advances in Neural Information Processing Systems, 2020, 33: 7198-7211.

[26]

LIU Y H, DE NADAI M, CAI D, et al. Describe what to change: a text-guided unsupervised image-to-image translation approach[C]// Proceedings of the 28th ACM International Conference on Multimedia. Seattle: ACM, 2020: 1357-1365.

[27]

KIM G, YE J C. DiffusionCLIP: text-guided image manipulation using diffusion models[EB/OL]. [2024-09-02]. https://doi.org/10.48550/arXiv.2110.02711.

[28]

SCHNEIDER D, SAQUIB SARFRAZ M, ROITBERG A, et al. Pose-based contrastive learning for domain agnostic activity representations[C]// 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 3432-3442.

[29]

SIAROHIN A, SANGINETO E, LATHUILIÈRE S, et al. Deformable GANs for pose-based human image generation[C]// 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 3408-3416.

[30]

LUNA-JIMÉNEZ C, KLEINLEIN R, GRIOL D, et al. A proposal for multimodal emotion recognition using aural transformers and action units on RAVDESS dataset[J]. Applied Sciences, 2022, 12(1): 327.

[31]

SAVCHENKO A V. HSEmotion: high-speed emotion recognition library[J]. Software Impacts, 2022, 14: 100433.

[32]

YI R, LIU Y J, LAI Y K, et al. APDrawingGAN: generating artistic portrait drawings from face photos with hierarchical GANs[C]// 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2020: 10735-10744.

[33]

LIU B C, SONG K P, ZHU Y Z, et al. Sketch-to-art: synthesizing stylized art images from sketches[C]// 2020 Asian Conference on Computer Vision. Cham: Springer, 2021: 207-222.

[34]

LIU X T, WU W L, LI C Z, et al. Reference-guided structure-aware deep sketch colorization for cartoons[J]. Computational Visual Media, 2022, 8(1): 135-148.

[35]

YUAN M C, SIMO-SERRA E. Linear art colorization with concatenated spatial attention[C]// 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021: 3941-3945.

[36]

WOO S, PARK J, LEE J Y, et al. CBAM: convolutional block attention module[C]// 2018 European Conference on Computer Vision. Cham: Springer, 2018: 3-19.

[37]

YU T, LI X, CAI Y F, et al. S2-MLPv2: improved spatial-shift MLP architecture for vision [EB/OL]. [2024-09-02]. https://doi.org/10.48550/arXiv.2108.01072.

[38]

HUANG X, BELONGIE S. Arbitrary style transfer in real-time with adaptive instance normalization[C]// 2017 IEEE International Conference on Computer Vision. Venice: IEEE, 2017: 1510-1519.

[39]

IIZUKA S, SIMO-SERRA E. DeepRemaster: temporal source-reference attention networks for comprehensive video enhancement[J]. ACM Transactions on Graphics, 2019, 38(6): 1-13.

[40]

黄小芬, 林丽群, 卢宇. 注意力残差密集网络的单幅图像去雾算法[J]. 福建师范大学学报(自然科学版), 2023, 39(1): 68-74.

[41]

XIE S N, TU Z W. Holistically-nested edge detection[J]. International Journal of Computer Vision, 2017, 125(1): 3-18.

[42]

WANG Z, BOVIK A C, SHEIKH H R, et al. Image quality assessment: from error visibility to structural similarity[J]. IEEE Transactions on Image Processing, 2004, 13(4): 600-612.

[43]

ZHANG R, ISOLA P, EFROS A A, et al. The unreasonable effectiveness of deep features as a perceptual metric[C]// 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 586-595.

[44]

GRIGORESCU C, PETKOV N, WESTENBERG M A. Contour detection based on nonclassical receptive field inhibition[J]. IEEE Transactions on Image Processing, 2003, 12(7): 729-739.

[45]

HUANG X, LIU M Y, BELONGIE S, et al. Multimodal unsupervised image-to-image translation[C]// 2018 European Conference on Computer Vision. Cham: Springer, 2018: 179-196.

[46]

ASHTARI A, SEO C W, KANG C, et al. Reference based sketch extraction via attention mechanism[J]. ACM Transactions on Graphics, 2022, 41(6): 1-16.

AI Summary AI Mindmap
PDF (2111KB)

44

访问

0

被引

详细

导航
相关文章

AI思维导图

/