In the backdrop of today's big data milieu on social media and online platforms, image-text sentiment analysis has become an important research task, which is crucial for understanding users' emotional tendencies. Existing methods are usually limited to the single-level features of modalities, lack an understanding of multi-level emotional information in images, and are prone to information redundancy and feature shift during multimodal feature fusion, resulting in poor model performance. In response to these issues, this study proposes an image-text sentiment analysis based on semantic guided attention and multi-task learning. Capture multilevel emotional information of images through a multiscale feature extraction module, use semantic-guided attention to fuse image information related to textual emotional information, and introduce an emotional focus calibration task in the multi-task learning module to minimize the distance between the fused features and their emotional centroids. Experimental results obtained from three social media datasets demonstrate that the proposed method outperforms existing methods in image and text sentiment analysis tasks.
WANGJ H, LIUZ, LIUT T,et al. Multimodal sentiment analysis based on multilevel feature fusion attention network[J]. Journal of Chinese Information Processing.2022, 36(10):145-154 (Ch).
SONGY F, RENG, YANGY,et al. Mutimodal sentiment analysis based on hybrid feature fusion of multi-level attention mechanism and multitask learning[J]. Application Research of Computers, 2022, 39(3):716-720. DOI: 10.19734/j.issn.1001-3695.2021.08.0357(Ch ).
HUJ, LIUY, ZHAOJ, et al. MMGCN: Multimodal fusion via deep graph convolution network for emotion recognition in conversation[EB/OL]. [2024-03-04]. DOI: 10.48550/arXiv.2107.06779 .
[8]
TSAIY H, BAIS J, LIANGP P, et al. Multimodal transformer for unaligned multimodal language sequences[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: Association for Computational Linguistics, 2019: 6558-6569. DOI: 10.18653/v1/p19-1656 .
[9]
HANW, CHENH, GELBUKHA, et al. Bi-bimodal modality fusion for correlation-controlled multimodal sentiment analysis[C]//Proceedings of the 23rd International Conference on Multimodal Interaction. New York: ACM Press,2021: 6-15. DOI: 10.1145/3462244.3479919 .
[10]
HUD, HOUX L, WEIL W, et al. MM-DFN: Multimodal dynamic fusion network for emotion recognition in conversations[C]// Proceedings of the 47th IEEE International Conference on Acoustics, Speech and Signal Processing. Piscataway: IEEE Press,2022: 7037-7041. DOI: 10.1109/ICASSP43922.2022.9747397 .
[11]
HER D, LEEW S, NGH T,et al. An interactive multi⁃task learning network for end-to-end aspect-based sentiment analysis[C]// Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: Association for Computational Linguistics, 2019:504⁃ 515. DOI: 10.18653/v1/p19-1048 .
[12]
AKHTARM S, CHAUHAND S, GHOSALD, et al. Multi-task learning for multi-modal emotion recognition and sentiment analysis[C]//Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies. Stroudsburg:Association for Computational Linguistics,2019:370-379. DOI: 10.18653/v1/n19-1034 .
[13]
JINN, WUJ X, MAX, et al. Multi-task learning model based on multi-scale CNN and LSTM for sentiment classification[J]. IEEE Access, 2020, 8: 77060-77072. DOI: 10.1109/access.2020.2989428 .
[14]
YANGB, WUL J, ZHUJ H, et al. Multimodal sentiment analysis with two-phase multi-task learning[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2022, 30: 2015-2024. DOI: 10.1109/TASLP.2022.3178204 .
[15]
MAJUMDERN, PORIAS, PENGH Y, et al. Sentiment and sarcasm classification with multitask learning[J]. IEEE Intelligent Systems, 2019, 34(3): 38-43. DOI: 10.1109/MIS.2019.2904691 .
[16]
YANGL, NAJ C, YUJ F. Cross-modal multitask transformer for end-to-end multimodal aspect-based sentiment analysis[J]. Information Processing & Management, 2022, 59(5): 103038. DOI: 10.1016/j.ipm.2022.103038 .
[17]
LIJ, ZHAOH W. Multimodal sentiment analysis method based on multi-task learning[C]//Proceedings of the 2023 6th International Conference on Signal Processing and Machine Learning. New York: ACM Press, 2023: 308-314. DOI: 10.1145/3614008.3614055 .
[18]
YUW M, XUH, YUANZ, et al. Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis[C]//Proceedings of the 35th AAAI Conference on Artificial Intelligence. Palo Alto:AAAI Press,2021: 10790-10797. DOI: 10.1609/aaai.v35i12.17289 .
[19]
WEIY W, YUANS Z, YANGR S, et al. Tackling modality heterogeneity with multi-view calibration network for multimodal sentiment detection[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: Association for Computational Linguistics, 2023: 5240-5252. DOI: 10.18653/v1/2023.acl-long.287 .
[20]
LIZ, XUB, ZHUC H, et al. CLMLF: A contrastive learning and multi-layer fusion method for multimodal sentiment detection[C]//Proceedings of the 19th North American Chapter of the Association for Computational Linguistics. Stroudsburg: Association for Computational Linguistics,2022:2282-2294. DOI: 10.18653/v1/2022.findings-naacl.175 .
[21]
HAZARIKAD, ZIMMERMANNR, PORIAS. MISA: Modality-invariant and-specific representations for multimodal sentiment analysis[C]//Proceedings of the 28th ACM International Conference on Multimedia. New York: ACM Press, 2020:1122-1131. DOI: 10.1145/3394171.3413678 .
[22]
LIUY, OTT M, GOYALN, et al. RoBERTa: A robustly optimized BERT pretraining approach[EB/OL]. [2024-03-04].
[23]
ZHAIP H, ZHANGD Y. Bidirectional-GRU based on attention mechanism for aspect-level sentiment analysis[C]//Proceedings of the 11th International Conference on Machine Learning and Computing. New York: ACM Press,2019: 86-90. DOI: 10.1145/3318299.3318368 .
[24]
DOSOVITSKIYA, BEYERL, KOLESNIKOVA, et al. An image is worth 16×16 words: Transformers for image recognition at scale [EB/OL]. [2024-07-21].
CAOM L.Sentiment analysis for images-text on social media based on auxiliary information extraction and fusion[D]. Wuhan: Wuhan University of Science and Technology,2022. DOI: 10.27380/d.cnki.gwkju.2022.000499(Ch ).
[27]
GEF, LIW, RENH, et al. Towards exploiting sticker for multimodal sentiment analysis in social media: A new dataset and baseline[C]//Proceedings of the 29th International Conference on Computational Linguistics. New York: ACM Press, 2022: 6795-6804.
LIS X, HUH J, LIUM F. Sentiment analysis of social media images-text based on semantic sense consistency[J]. China Sciencepaper, 2023, 18(3): 322-329 (Ch).
[30]
KENTONJ D M W C, TOUTABOVAL K. Bert: Pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the 17th NAACL-HLT. Stroudsburg: Association for Computational Linguistics, 2019:4171-4186.
[31]
HEK M, ZHANGX Y, RENS Q, et al. Deep residual learning for image recognition[C]// Proceedings of the 22nd IEEE Winter Conference on Applications of Computer Vision. Piscataway: IEEE Press, 2016: 770-778. DOI: 10.1109/CVPR.2016.90 .
[32]
YUJ F, JIANGJ. Adapting BERT for target-oriented multimodal sentiment classification[C]//Proceedings of the 28th International Joint Conference on Artificial Intelligence. Freiburg: Morgan Kaufmann,2019:5408-5414. DOI: 10.24963/ijcai.2019/751 .
[33]
TRUONGQ T, LAUWH W. VistaNet: Visual aspect attention network for multimodal sentiment analysis[C]// Proceedings of the 33th AAAI Conference on Artificial Intelligence. Palo Alto:AAAI Press,2019: 305-312. DOI: 10.1609/aaai.v33i01.3301305 .
HUH J, DINGZ Y, ZHANGY F, et al. Images-text sentiment analysis in social media based on joint and interactive attention[J/OL]. Journal of Beijing University of Aeronautics and Astronautics, 1-11. [2023-11-23].