Collaborative Perception: Multi-Task Driven Adaptive Multimodal Sentiment Analysis

CAO Hong-fei, CAO Tao

Computer & Telecommunication ›› 2026 ›› Issue (2) : 29-34.

Computer & Telecommunication ›› 2026 ›› Issue (2) : 29-34.

Collaborative Perception: Multi-Task Driven Adaptive Multimodal Sentiment Analysis

  • CAO Hong-fei, CAO Tao
Author information +
History +

Abstract

To address the issues of feature redundancy and noise interference caused by modal heterogeneity in multimodal sentiment analysis, this paper proposes an adaptive collaborative perception model based on multi-task joint learning (AdaSP-MTL). The AdaSP-MTL framework decouples unimodal and multimodal sentiment tasks and constructs a heterogeneous feature disentanglement space to enhance modality-specific modeling. A dynamic gated fusion module based on inter-modal cosine similarity is designed to adaptively allocate fusion weights, thereby suppressing cross-modal noise propagation. Furthermore, the model integrates low-level local features extracted by CNNs with high-level sequential features captured by BiGRUs to form multi-scale sentiment representations. Experimental evaluations on the SIMS, MOSI, and MOSEI datasets demonstrate that AdaSP-MTL comprehensively outperforms existing baseline models in sentiment classification tasks, validating its advantages in effectively modeling modal heterogeneity, suppressing noise interference, and constructing robust sentiment representations.

Key words

multimodal sentiment analysis / multi-task learning / heterogeneous features / adaptive fusion / multi-scale features / multimodal fusion

Cite this article

Download Citations
CAO Hong-fei, CAO Tao. Collaborative Perception: Multi-Task Driven Adaptive Multimodal Sentiment Analysis[J]. Computer & Telecommunication. 2026(2): 29-34

References

[1] Holler J,Levinson S C.Multimodal Language Processing in Human Communication[J].Trends in Cognitive Sciences,2019,23(8):639-652.
[2] Miao H,Zhang Y,Wang D,et al.Multi-Output Learning Based on Multimodal GCN and Co-Attention for Image Aesthetics and Emotion Analysis[J].Mathematics,2021,9(12):1437.
[3] Singh P,Srivastava R,Rana K P S,et al.A Multimodal Hierarchical Approach to Speech Emotion Recognition from Audio and Text[J].Knowledge-Based Systems,2021,229:107316.
[4] Zhang Q,Wei Y,Han Z,et al.Multimodal Fusion on Low-Quality Data:A Comprehensive Survey[EB/OL].2024:arXiv:2404.18947.https://arxiv.org/abs/2404.18947
[5] Zhang S,Yang Y,Chen C,et al.Deep Learning-Based Multimodal Emotion Recognition from Audio,Visual,and Text Modalities:A Systematic Review of Recent Advancements and Future Prospects[J].Expert Systems with Applications,2024,237:121692.
[6] Devlin J,Chang M W,Lee K,et al.BERT:Pre-Training of Deep Bidirectional Transformers for Language Understanding[EB/OL].2018:arXiv:1810.04805. https://arxiv.org/abs/1810.04805.
[7] Xu N,Mao W,Chen G.A Co-Memory Network for Multimodal Sentiment Analysis[C]//The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval.ACM,2018:929-932.
[8] Mai S,Hu H,Xing S.Divide,Conquer and Combine:Hierarchical Feature Fusion Network with Local and Global Perspectives for Multimodal Affective Computing[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.Stroudsburg,PA,USA:ACL,2019:481-492.
[9] Zadeh A,Liang P P,Poria S,et al.Multi-Attention Recurrent Network for Human Communication Comprehension[J].Proceedings of the AAAI Conference on Artificial Intelligence,2018,32(1):5642-5649.
[10] Mai S,Hu H,Xu J,et al.Multi-Fusion Residual Memory Network for Multimodal Human Sentiment Comprehension[J].IEEE Transactions on Affective Computing,2022,13(1):320-334.
[11] McFee B,Raffel C,Liang D,et al.Librosa:Audio and Music Signal Analysis in Python[J].Proceedings of the 14th Python in Science Conference,2015:18-24.
[12] Baltrušaitis T,Robinson P,Morency L P.OpenFace:An Open Source Facial Behavior Analysis Toolkit[C]//2016 IEEE Winter Conference on Applications of Computer Vision (WACV).IEEE,2016:1-10.
[13] Zadeh A,Zellers R,Pincus E,et al.MOSI:Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos[EB/OL].2016:arXiv:1606.06259. https://arxiv.org/abs/1606.06259.
[14] Bagher Zadeh A,Liang P P,Poria S,et al.Multimodal Language Analysis in the Wild:CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph[C]//Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1:Long Papers).Association for Computational Linguistics,2018:2236-2246.
[15] Yu W,Xu H,Meng F,et al.CH-SIMS:A Chinese Multimodal Sentiment Analysis Dataset with Fine-Grained Annotation of Modality[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.Association for Computational Linguistics,2020:3718-3727.
[16] Zadeh A,Chen M,Poria S,et al.Tensor Fusion Network for Multimodal Sentiment Analysis[EB/OL].2017:arXiv:1707.07250.https://arxiv.org/abs/1707.07250.
[17] Liu Z,Shen Y,Lakshminarasimhan V B,et al.Efficient Low-Rank Multimodal Fusion with Modality-Specific Factors[EB/OL].2018:arXiv:1806.00064. https://arxiv.org/abs/1806.00064.
[18] A.Zadeh,P.P.Liang,N.Mazumder,S.Poria,E.Cambria,L.-P.Morency,"Memory fusion network for multi-view sequential learning," in Proc.AAAI Conf.Artif.Intell.,vol.32,no.1,Feb.2018.
[19] Koromilas P,Nicolaou M A,Giannakopoulos T,et al.MMATR:A Lightweight Approach for Multimodal Sentiment Analysis Based on Tensor Methods[C]//ICASSP 2023 - 2023 IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP).IEEE,2023:1-5.
[20] Peng J,Wu T,Zhang W,et al.A Fine-Grained Modal Label-Based Multi-Stage Network for Multimodal Sentiment Analysis[J].Expert Systems with Applications,2023,221:119721.
[21] Wang L,Peng J,Zheng C,et al.A Cross Modal Hierarchical Fusion Multimodal Sentiment Analysis Method Based on Multi-Task Learning[J].Information Processing & Management,2024,61(3):103675.
[22] 陈巧红,孙佳锦,漏杨波,等.基于多任务学习与层叠Transformer的多模态情感分析模型[J].浙江大学学报(工学版),2023,57(12):2421-2429.
[23] 王有康,程春玲.基于跨模态单向加权的多模态情感分析模型[J].计算机科学,2025,52(7):226-232.

Accesses

Citation

Detail

Sections
Recommended

/