In the existing research on multi-modal sentiment analysis, the fusion different modal information is mainly through the overall interaction of different modal features, but it doesn’t consider the relationship between unique features and common features contained in different modes, so the complex emotions can’t be analyzed effectively. To solve this problem, a multimodal sentiment analysis model based on adaptive gated decoupling feature fusion (AGDF) was proposed. Firstly, the pre-trained BERT model and Transformer model were used for feature extraction of different modes. Secondly, according to the principle that the common features of different modes were similar but the unique features were not similar, the contrast pair was constructed. By contrastive learning, the features of different modes were decomposed into unique features and common features. Thirdly, according to the principle that the image and speech modes were offset in the text mode, a new adaptive gating mechanism was designed to fuse the features and integrate other modal information into the text mode. At the same time, the relation graph of unique features and common features was designed, and the fusion of the graph attention neural network was used to balance the unique information and common information among the modes. Finally, the fusion features were classified. Experiments on the datasets CMU-MOSI and CMU-MOSEI show that the accuracy and F1 score of the proposed method are improved by about 1 percentage point compared with the baseline method. In addition, compared with other feature decomposition methods, the proposed method improves accuracy by 1.23 percentage point, F1 score by 1.37 percentage point, Corr by 2.13 percentage point, and reduces MAE by 4.83 percentage point. Consequently, the proposed method can make full use of the heterogeneous information of different modes and effectively improve the effect of sentiment analysis.
LIUN, ZHAOJ. Recommendation system based on deep sentiment analysis and matrix factorization[J]. IEEE Access, 2023, 11: 16994-17001.
[2]
PARKJ, SEOY S. Twitter sentiment analysis-based adjustment of cryptocurrency action recommendation model for profit maximization[J]. IEEE Access, 2023, 11: 44828-44841.
[3]
WANGJ, CHENZ. SPCM: A machine learning approach for sentiment-based stock recommendation system[J]. IEEE Access, 2024, 12: 14116-14129.
[4]
LIQ, GKOUMASD, LIOMAC, et al. Quantum-inspired multimodal fusion for video sentiment analysis[J]. Information Fusion, 2021, 65: 58-71.
[5]
ZHANGY, SONGD, LIX, et al. A quantum-like multimodal network framework for modeling interaction dynamics in multiparty conversational sentiment analysis[J]. Information Fusion, 2020, 62: 14-31.
[6]
WANGF, TIANS, YUL, et al. TEDT: Transformer-based encoding–decoding translation network for multimodal sentiment analysis[J]. Cognitive Computation, 2022, 15(1): 289-303.
HAZARIKAD, ZIMMERMANNR, PORIAS. MISA: Modality-invariant and -specific representations for multimodal sentiment analysis[C]//Proceedings of the 28th ACM International Conference on Multimedia. New York, NY, USA, 2020: 1122-1131.
[9]
VAN AMSTERDAMB, KADKHODAMOHAMMA- DIA, LUENGOI, et al. ASPnet: Action segmentation with shared-private representation of multiple data sources[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023: 2384-2393.
[10]
LUOY, WUR, LIUJ, et al. Balanced sentimental information via multimodal interaction model[J]. Multimedia Systems, 2024, 30(1): 10.
[11]
YANGJ, YUY, NIUD, et al. ConFEDE: Contrastive feature decomposition for multimodal sentiment analysis[C]//Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Toronto, Canada, 2023: 7617-7630.
[12]
HUANGJ, TAOJ, LIUB, et al. Multimodal transformer fusion for continuous emotion recognition[C]//ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020: 3507-3511.
[13]
PRAVEENR G, GRANGERE, CARDINALP. Cross attentional audio-visual fusion for dimensional emotion recognition[C]//2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021), 2021: 1-8.
[14]
FARHOUDIZ, SETAYESHIS. Fusion of deep learning features with mixture of brain emotional learning for audio-visual emotion recognition[J]. Speech Communication, 2021, 127: 92-103.
[15]
GKOUMASD, LIQ, LIOMAC, et al. What makes the difference? An empirical comparison of fusion strategies for multimodal language analysis[J]. Information Fusion, 2021, 66: 184-197.
[16]
TANGJ, LIUD, JINX, et al. BAFN: Bi-direction attention based fusion network for multimodal sentiment analysis[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 33(4): 1966-1978.
[17]
WUT, PENGJ, ZHANGW, et al. Video sentiment analysis with bimodal information-augmented multi-head attention[J]. Knowledge-Based Systems, 2022, 235: 107676.
[18]
LIZ, GUOQ, PANY, et al. Multi-level correlation mining framework with self-supervised label generation for multimodal sentiment analysis[J]. Information Fusion, 2023, 99: 101891.
[19]
CHENT, KORNBLITHS, NOROUZIM, et al. A simple framework for contrastive learning of visual representations[DB/OL].(2020-03-30)[2024-07-05].
[20]
CHENX, HEK. Exploring Simple siamese representation learning[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021: 15745-15753.
[21]
RADFORDA, KIMJ W, HALLACYC, et al. Learning transferable visual models from natural language supervision[C]//Proceedings of the 38th International Conference on Machine Learning, 2021: 8748-8763.
[22]
LIX, SUNA, ZHAOM, et al. Multi-intention oriented contrastive learning for sequential recommendation[C]//Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining. New York, NY, USA, 2023: 411-419.
[23]
BRODYS, ALONU, YAHAVE. How attentive are graph attention networks?[DB/OL].(2022-01-31)[2024-07-05].
[24]
YOUJ, DUT, LESKOVECJ. ROLAND: Graph learning framework for dynamic graphs[DB/OL].(2022-08-15)[2024-07-05].
[25]
LIK, HUANGZ, JIAZ. RAHG: A role-aware hypergraph neural network for node classification in graphs[J]. IEEE Transactions on Network Science and Engineering, 2023, 10(4): 2098-2108.
[26]
TANGZ, XIAOQ, QINY, et al. Multi-view interactive representations for multimodal sentiment analysis[J]. IEEE Transactions on Consumer Electronics, 2024, 70(1): 4095-4107.
[27]
RAHMANW, HASANM K, LEES, et al. Integrating Multimodal Information in Large Pretrained Transformers[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020: 2359-2369.
[28]
ZADEHA, ZELLERSR, PINCUSE, et al. MOSI: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos[DB/OL].(2016-08-12)[2024-07-05].
[29]
BAGHER ZADEHA, LIANGP P, PORIAS, et al. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph[C]//Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. Melbourne, Australia, 2018: 2236-2246.
[30]
TSAIY H H, BAIS, LIANGP P, et al. Multimodal transformer for unaligned multimodal language sequences[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy, 2019: 6558-6569.
[31]
YUW, XUH, YUANZ, et al. Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis[DB/OL].(2021-02-09)[2024-07-05].
[32]
HANW, CHENH, PORIAS. Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis[C]//Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021: 9180-9192.
[33]
WUJ, MAI S, HUH. Graph capsule aggregation for unaligned multimodal sequences[C]//Proceedings of the 2021 International Conference on Multimodal Interaction. New York, NY, USA, 2021: 521-529.
[34]
XUM, LIANGF, SUX, et al. CMJRT: Cross-modal joint representation transformer for multimodal sentiment analysis[J]. IEEE Access, 2022, 10: 131671-131679.