1.School of Computer Science and Technology,Wuhan University of Science and Technology,Wuhan 430065,Hubei,China
2.Hubei Province Key Laboratory of Intelligent Information Processing and Real-Time Industrial System,Wuhan 430065,Hubei,China
3.Big Data Science and Engineering Research Institute,Wuhan University of Science and Technology,Wuhan 430065,Hubei,China
Show less
文章历史+
Received
Published
2022-08-26
2023-04-24
Issue Date
2026-07-23
PDF (3268K)
摘要
在传统的图像描述生成任务中,已有方法对图像的描述仅仅停留在浅层,并缺乏真实世界知识的指导,难以挖掘出对象在特定背景下的逻辑语义关系。新闻文本的引入为图像描述带来了新的可能,同时对模型的学习能力有了更高要求;此外,新闻图集中往往存在多幅图像,且相互之间联系紧密,导致现有单图描述生成方法不适用于新闻图集描述生成。针对上述问题,本文提出了一种基于图文双向引导注意力(image and text bidirectional guidance attention,ITBGA)的新闻图集描述方法,以图集作为研究对象,并辅以对应的新闻文本作为背景知识,基于ITBGA分别实现粗、细两个粒度的跨模态信息交互,并通过指针网络辅助命名实体词生成。在本文构建的新闻图集数据集上进行了实验验证,结果表明ITBGA能有效提升描述文本的质量,在关键的CIDEr指标上达到了最优。
Abstract
In the traditional image captioning task, each method only describes the image in a shallow level. Due to the lack of real world knowledge, it is often difficult to mine the logical semantic relationship of objects in a specific context. The introduction of news text brings new possibilities for image captioning, but at the same time, it requires higher learning ability of models; In addition, there are often multiple images in news data, and they are closely related to each other, which makes the existing single image captioning methods not suitable for news image set captioning task. To solve the above problems, this paper proposes a news image set captioning method based on image and text bidirectional guidance attention, i.e., ITBGA, which takes the image set as the research object and the corresponding news text as the background knowledge. Based on ITBGA, it realizes the cross modal information interaction at coarse and fine granularity respectively, and uses the pointer network to assist the generation of named entity words. The experimental results on the news image set constructed in this paper show that ITBGA can improve the quality of the description text, and the method achieves optimal performance on CIDEr indicator.
针对上述分析,本文以国内新闻图集网站作为数据源,构建了新闻图集描述数据集,其中每条数据包含多幅图像和一段背景文本,对应一条人工书写的描述文本。在此基础上,本文提出了新闻图集描述生成方法(news image set captioning,NISC),利用背景信息为图集生成解释性的长描述文本。在该方法中引入了图文双向引导注意力(image and text bidirectional guidance attention,ITBGA),针对新闻图集的特性,设计了两种不同的形态:粗粒度图文双向引导注意力(coarse-grained image text bidirectional guidance attention,CITBGA)和细粒度图文双向引导注意力(fine-grained image text bidirectional guidance attention,FITBGA),结合外部文本进行双向引导关注,既对图像和图像内容进行选择,也对文本信息进行选择,使得模型能够充分利用图文信息,并有效筛选出关键信息;同时将得到的注意力权重用于指针网络,引导命名实体生成。
GN-WAC(word2vec+avg+cintx):原型是由Biten等[15]提出的一种基于模板的方法,描述过程分为两个步骤:生成模板,插入命名实体。生成模型整体以show attend and tell[5]作为基础结构。对文章的编码以句子为单位,先通过word2vec方式获得词向量,再对词向量求平均值得到句子的向量表示。图文特征之间仅拼接,不进行额外交互。生成句子模板后,第二步使用context insertion方式,对具有余弦相似性的文章句子进行排序,实现命名实体插入。
MAQ X, LIP J, SONGJ Y, et al. The development trends and applications of image caption[J]. Unmanned Systems Technology, 2020, 3(6): 25-35. DOI: 10.19942/j.issn.2096-5915.2020.06.003(Ch ).
[3]
CHOK, VAN MERRIENBOERB, GULCEHREC, et al. Learning phrase representations using RNN encoder-decoder for statistical machine translation[EB/OL]. [2022-05-14]. DOI: 10.3115/v1/d14-1179 .
[4]
VINYALSO, TOSHEVA, BENGIOS, et al. Show and tell: A neural image caption generator[C]//2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2015: 3156-3164. DOI: 10.1109/CVPR.2015.7298935 .
[5]
VASWANIA, SHAZEERN, PARMARN, et al. Attention is all you need[C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. New York: ACM, 2017: 6000-6010. DOI: 10.5555/3295222.3295349 .
[6]
XUK, BAJ L, KIROSR, et al. Show, attend and tell: Neural image caption generation with visual attention[C]//Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37. New York: ACM, 2015: 2048-2057. DOI: 10.5555/3045118.3045336 .
[7]
ANDERSONP, HEX D, BUEHLERC, et al. Bottom-up and top-down attention for image captioning and visual question answering[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 6077-6086. DOI: 10.1109/CVPR.2018.00636 .
[8]
LIN, CHENZ. Image Cationing with Visual-Semantic LSTM[C]//Proceedings of the 27th International Joint Conference on Artificial Intelligence. Stockholm: AAAI Press, 2018: 793-799. DOI: 10.24963/ijcai.2018/110 .
LIUM F, BIJ Q, ZHOUB Y, et al. Interpretable image caption generation based on dependency syntax[J/OL]. Journal of Computer Research and Development, 2022: 1-12. (2022-10-27).
SHIY L, YANGW Z, DUH X, et al. Overview of image captions based on deep learning[J]. Acta Electronica Sinica, 2021, 49(10): 2048-2060. DOI: 10.12263/DZXB.20200669(Ch ).
[15]
ZHAOS Q, SHARMAP, LEVINBOIMT, et al. Informative image captioning with external sources of information[EB/OL]. 2019: arXiv: 1906.08876. DOI: 10.18653/v1/p19-1650 .
[16]
RAMISAA, YANF, MORENO-NOGUERF, et al. BreakingNews: Article annotation by image and text processing[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 40(5): 1072-1085. DOI: 10.1109/TPAMI.2017.2721945 .
[17]
TRANA, MATHEWSA, XIEL X. Transform and tell: Entity-aware news image captioning[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 13032-13042. DOI: 10.1109/CVPR42600.2020.01305 .
BITENA F, GOMEZL, RUSIÑOLM, et al. Good news, everyone! context driven entity-aware captioning for news images[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 12458-12467. DOI: 10.1109/CVPR.2019.01275 .
[20]
LIUF X, WANGY H, WANGT L, et al. Visual news: Benchmark and challenges in news image captioning[EB/OL]. 2020: arXiv: 2010.03743. DOI: 10.18653/v1/2021.emnlp-main.542 .
[21]
SEE A, LIUP J, MANNINGC D. Get to the point: Summarization with pointer-generator networks[EB/OL]. 2017: arXiv: 1704.04368. DOI: 10.18653/v1/p17-1099 .
[22]
WANGH X, KAWAHARAY, WENGC Q, et al. Representative selection with structured sparsity[J]. Pattern Recognition, 2017, 63: 268-278. DOI: 10.1016/j.patcog.2016.10.014 .
[23]
SIMONI, SNAVELYN, SEITZS M. Scene summarization for online image collections[C]//2007 IEEE 11th International Conference on Computer Vision. New York: IEEE Press, 2007: 1-8. DOI: 10.1109/ICCV.2007.4408863 .
[24]
YANGC L, SHENJ L, PENGJ Y, et al. Image collection summarization via dictionary learning for sparse representation[J]. Pattern Recognition, 2013, 46(3): 948-961. DOI: 10.1016/j.patcog.2012.07.011 .
[25]
TSCHIATSCHEKS, IYERR, WEIH C, et al. Learning mixtures of submodular functions for image collection summarization[C]//Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 1. New York: ACM, 2014: 1413-1421.
[26]
SINGLAA, TSCHIATSCHEKS, KRAUSEA. Noisy submodular maximization via adaptive sampling with applications to crowdsourced image collection summarization[C]//Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence. New York: ACM, 2016: 2037-2043. DOI: 10.5555/3016100.3016183 .
[27]
ABDU JYOTHIA. Generating Natural Language Summary for Image Sets[D]. Burnaby: Simon Fraser University, 2018.
[28]
PLUMMERB A, WANGL W, CERVANTESC M, et al. Flickr30k entities: collecting region-to-phrase correspondences for richer image-to-sentence models[C]//2015 IEEE International Conference on Computer Vision (ICCV). New York: IEEE Press, 2016: 2641-2649. DOI: 10.1109/ICCV.2015.303 .
[29]
CHENX L, FANGH, LINT Y, et al. Microsoft COCO captions: Data collection and evaluation server[EB/OL]. 2015: arXiv: 1504.00325.
[30]
HEK M, ZHANGX Y, RENS Q, et al. Deep residual learning for image recognition[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2016: 770-778. DOI: 10.1109/CVPR.2016.90 .
[31]
JIAOZ Y, SUNS Q, SUNK. Chinese lexical analysis with deep Bi-GRU-CRF network[EB/OL]. 2018: arXiv: 1807.01882.
[32]
DEVLINJ, CHANGM W, LEEK, et al. BERT: Pre-training of deep bidirectional transformers for language understanding[EB/OL]. 2018: arXiv: 1810.04805. DOI: 10.48550/arXiv.1810.04805 .
[33]
MIKOLOVT, CHENK, CORRADOG, et al. Efficient estimation of word representations in vector space[EB/OL]. 2013: arXiv: 1301.3781.
[34]
PAPINENIK, ROUKOSS, WARDT, et al. BLEU: A method for automatic evaluation of machine translation[C]//Proceedings of the 40th Annual Meeting on Association for Computational Linguistics. New York: ACM, 2002: 311-318. DOI: 10.3115/1073083.1073135 .
[35]
LINC Y. ROUGE: A package for automatic evaluation of summaries[DB/OL].[2022-08-06]. DOI: 10.3115/1218955.1219032 .
[36]
BANERJEES, LAVIEA. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments[C]//Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization. Stroudsburg: Association for Computational Linguistics, 2005: 65-72.
[37]
VEDANTAMR, LAWRENCE ZITNICKC, PARIKHD. CIDEr: Consensus-based image description evaluation[C]//2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2015: 4566-4575. DOI: 10.1109/CVPR.2015.7299087 .