Dynamic virtual try-on aims to generate coherent, smooth, and realistic fitting videos. Current methods often encounter issues such as clothing self-occlusion and blurry patterns due to changes in body posture.Therefore, this paper proposes a clothing distortion constraint and prediction method based on the Spatial Transformer Network (STN). In the clothing distortion network, the Transformer module takes advantage of both global information and local key information to strengthen the data feature region, and the STN module uses the learnable Thin Plate Spline interpolation (TPS) method to predict the clothing distortion range and obtain the distorted image and mask. The try-on network is conducive to the U-Net network of self-attention mechanisms to align the distorted image mask and human body representation information, and generate high-quality try-on images. Finally, a dynamic synthesis network was used to solve the temporal consistency problem of video frames, and a coherent high-quality fitting video was generated. On the VVT dataset, compared to CP-VTON, our proposed method achieved an improvement of 0.076 in the average Structural Similarity Index (SSIM) and a decrease of 0.420 in the average perceptual image patch similarity (LPIPS). Compared to the FW-GAN method, it reduced by 0.089 in the I3D metric and 2.252 in the ResNeXt101 metric. On the VITON-HD dataset, the SSIM index of the proposed method exceeds that of CP-VTON and FW-GAN, further indicating that the images generated by the proposed method exhibit high quality and low distortion.
在动态合成网络中,需要同时考虑当前帧和过去的帧。首先,从这些帧中通过有限元分析(Finite Element Analysis,FEA)计算获取光流信息。再利用光流信息的方向注释,通过1.1节中的服装扭曲网络将之前生成的帧扭曲(或调整)到当前时间段的位置。最后,使用一个掩码将扭曲后的帧与原始合成的结果融合,以生成最终的合成帧。接下来将等中间输出进行合成,最后生成试衣视频。
GOODFELLOWI, POUGET-ABADIEJ, MIRZAM, et al. Generative adversarial nets[C]//Advances in Neural Information Processing Systems 27 (NIPS 2014). Cambridge: MIT Press, 2014:2672-2680.
[2]
GEC J, SONGY B, GEY Y, et al. Disentangled cycle consistency for highly-realistic virtual try-on[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2021: 16923-16932. DOI: 10.1109/CVPR46437.2021.01665 .
[3]
HANX T, HUX J, HUANGW L, et al. ClothFlow: A flow-based model for clothed person generation[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE Press, 2019: 10470-10479. DOI: 10.1109/ICCV.2019.01057 .
[4]
ISSENHUTHT, MARYJ, CALAUZÈNESC. Do not mask what you do not need to mask: A parser-free virtual try-on[C]//European Conference on Computer Vision. Cham: Springer, 2020: 619-635.10.1007/978-3-030-58565-5_37. DOI: 10.1007/978-3-030-58565-5_37 .
[5]
HANX T, WUZ X, WUZ, et al. VITON: An image-based virtual try-on network[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 7543-7552. DOI: 10.1109/CVPR.2018.00787 .
MATIURR M, THAIT T, HEEJUNEA, et al. CP-VTON+: Clothing shape and texture preserving image-based virtual try-on[DB/OL].[2023-02-01].
[8]
YANGH, ZHANGR M, GUOX B, et al. Towards photo-realistic virtual try-on by adaptively generating-preserving image content[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 7850-7859. DOI: 10.1109/CVPR42600.2020.00787 .
[9]
CHOIS W, PARKS Y, LEEM S, et al. VITON-HD: High-resolution virtual try-on via misalignment-aware normalization[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2021: 14131–14140. DOI: 10.1109/CVPR46437.2021.01391 .
[10]
DONGH Y, LIANGX D, SHENX H, et al. FW-GAN: Flow-navigated warping GAN for video virtual try-on[C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. New York: IEEE Press, 2019: 1161-1170. DOI: 10.1109/ICCV.2019.00125 .
KUPPAG, JONGA, LIUX, et al. ShineOn: illuminating design choices for practical video-based virtual clothing try-on[C]// Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. New York: IEEE Press, 2021: 191-200. DOI: 10.1109/WACVW52041.2021.00025 .
[13]
XIEL Z, HUANGZ Y, DONGX, et al. GP-VTON: Towards general purpose virtual try-on via collaborative local-flow global-parsing learning[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2023: 23550-23559. DOI: 10.1109/CVPR52729.2023.02255 .
[14]
BOOKSTEINF L. Principal warps: Thin-plate splines and the decomposition of deformations[C]//IEEE Transactions on Pattern Analysis and Machine Intelligence. New York: IEEE Press, 1989: 567-585. DOI: 10.1109/34.24792 .
[15]
BROXT, BRUHNA, PAPENBERGN, et al. High accuracy optical flow estimation based on a theory for warping[C]//European Conference on Computer Vision. Berlin: Springer, 2004: 25-36.10.1007/978-3-540-24673-2_3. DOI: 10.1007/978-3-540-24673-2_3 .
[16]
HUJ, SHENL, SUNG. Squeeze-and-excitation networks[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 7132-7141. DOI: 10.1109/CVPR.2018.00745 .
[17]
WANGX L, GIRSHICKR, GUPTAA, et al. Non-local neural networks[C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 7794-7803. DOI: 10.1109/CVPR.2018.00813 .
[18]
WANGZ, BOVIKA C, SHEIKHH R, et al. Image quality assessment: From error visibility to structural similarity[J]. IEEE Transactions on Image Processing: A Publication of the IEEE Signal Processing Society, 2004, 13(4): 600-612. DOI: 10.1109/tip.2003.819861 .
[19]
ZHANGR, ISOLAP, EFROSA A, et al. The unreasonable effectiveness of deep features as a perceptual metric[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 586-595. DOI: 10.1109/CVPR.2018.00068 .
[20]
CARREIRAJ, ZISSERMANA. Quo vadis, action recognition? A new model and the kinetics dataset[C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.New York: IEEE Press, 2017: 4724-4733. DOI: 10.1109/CVPR.2017.502 .
[21]
HARAK, KATAOKAH, SATOHY. Can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet [C]// Proceedings of the IEEE conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 6546-6555. DOI: 10.1109/CVPR.2018.00685 .
[22]
HES, SONGY Z, XIANGT. Style-based global appearance flow for virtual try-on[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2022: 3470-3479. DOI: 10.1109/CVPR52688.2022.00346 .