In response to the scarcity of resources and the challenges related to low quality in the Tibetan text-to-image field, a Tibetan Text-to-Image Generation Method Based on Transformer and Generative Adversarial Networks is proposed in this paper. The method uses the Transformer architecture to train text encoders with different granularity to extract Tibetan text features. After affine transformation, these text features are then fused with noise obtained from random sampling, and convolutional layers are used to generate images. Experimental results show that, on the CUB-BO dataset constructed in this paper, the Inception Score (IS) and Fréchet Inception Distance (FID) values reach 5.22 and 14.43, respectively, indicating a high capability of generating images from Tibetan text. In addition, it was found that images generated from Tibetan text processed using syllable segmentation performed better in terms of detail clarity and semantic consistency compared to those created using sub-word segmentation.
文本生成图像技术在处理英文和汉文等文字方面已取得显著进展,然而对于藏文等少数民族文字的处理仍待深入研究。藏文在文字结构、书写规则以及语义表达等方面与英文、汉文等文字存在显著差异,对文本生成图像模型的文本理解与图像生成能力带来了更大的挑战。此外,藏文生成图像数据集相对匮乏,数据资源的不足严重制约了模型的训练与优化,进而影响图像生成效果。同时,现有主流文本生成图像模型多基于英汉文等文字设计,在技术架构与算法实现上缺乏对少数民族文字的特殊性考虑,导致模型在处理藏文时面临技术适配性问题。为了推动藏文生成图像技术的发展,本文在DF-GAN模型基础上提出了基于Transformer[9]和生成对抗网络的藏文生成图像方法(A Text-to-Image Synthesis Method for the Tibetan Language Based on Transformer and Generative Adversarial Networks, TT-GAN)。本文的主要创新点如下:
(1) 通过机器翻译和人工校对的方式构建了藏文生成图像数据集CUB-BO。
(2) 在文本编码器中将原有的双向长短时记忆网络(Bi-directional Long Short-Term Memory, BiLSTM)[10]替换为Transformer。借助Transformer自注意力机制以更好地处理藏文文本中的长距离依赖关系。
GoodfellowI, Pouget-AbadieJ, MirzaM, et al. Generative adversarial networks[J]. Communications of the ACM, 2020, 63(11): 139-144
[2]
ReedS, AkataZ, YanX, et al. Generative adversarial text to image synthesis[C]. International Conference on Machine Learning. PMLR, 2016: 1060-1069.
[3]
ZhangH, XuT, LiH, et al. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks[C]. Proceedings of the IEEE International Conference on Computer Vision, 2017: 5907-5915.
[4]
ZhangH, XuT, LiH, et al. Stackgan++: Realistic image synthesis with stacked generative adversarial networks[J]. 2018 IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 41(8): 1947-1962.
[5]
XuT, ZhangP, HuangQ, et al. Attngan: Fine-grained text to image generation with attentional generative adversarial networks[C]. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018: 1316-1324.
[6]
TaoM, TangH, WuF, et al. Df-gan: A simple and effective baseline for text-to-image synthesis[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 16515-16525.
[7]
YeS, WangH, TanM, et al. Recurrent affine transformation for text-to-image synthesis[J]. IEEE Transactions on Multimedia, 2023, 26(000): 462-473.
VaswaniA. Attention is all you need[J]. Advances in Neural Information Processing Systems, 2017, 30: 5998-6008.
[10]
SchusterM, PaliwalK K. Bidirectional recurrent neural networks[J]. IEEE Transactions on Signal Processing, 1997, 45(11): 2673-2681.
[11]
SalimansT, GoodfellowI, ZarembaW, et al. Improved techniques for training GANs[C]. Proceedings of the 30th International Conference on Neural Information Processing Systems. Barcelona: Curran Associates Inc, 2016: 2234-2242.
[12]
HeuselM, RamsauerH, UnterthinerT, et al. GANs trained by a two time-scale update rule converge to a local nash equilibrium[C]. Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach: Curran Associates Inc, 2017: 6629-6640.
RadfordA, WuJ, ChildR, et al. Language models are unsupervised multitask learners[J]. OpenAI Blog, 2019, 1(8): 9.
[16]
DevlinJ, ChangM W, LeeK, et al. Bert: Pre-training of deep bidirectional transformers for language understanding[C]. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019: 4171-4186.
[17]
SennrichR, HaddowB, BirchA. Neural machine translation of rare words with subword units[C]. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2016: 1715-1725.
YangZ, XuZ, CuiY, et al. CINO: A Chinese minority pre-trained language model[C]. Proceedings of the 29th International Conference on Computational Linguistics. Gyeongju, Republic of Korea: International Committee on Computational Linguistics, 2022: 3937-3949.
KudoT, RichardsonJ. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing[C]. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,2018: 66-71.
[22]
SzegedyC, VanhouckeV, IoffeS, et al. Rethinking the inception architecture for computer vision[C]. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016: 2818-2826.
[23]
WAH C, BRANSONS, WELINDERP, et al. The Caltech-UCSD Birds-200-2011 dataset[EB/OL].(2022-08-12)[202405-26].
[24]
KingmaD P, BaJ. Adam: a method for stochastic optimization[EB/OL].(2014-12-22)[2023-03-18].
[25]
LiaoW, HuK, YangM Y, et al. Text to image generation with semantic-spatial aware gan[C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2022:18187-18196.