协同感知:多任务驱动的自适应多模态情感分析

曹鸿飞, 曹涛

电脑与电信 ›› 2026 ›› Issue (2) : 29-34.

电脑与电信 ›› 2026 ›› Issue (2) : 29-34.
智能感知与计算

协同感知:多任务驱动的自适应多模态情感分析

  • 曹鸿飞, 曹涛
作者信息 +

Collaborative Perception: Multi-Task Driven Adaptive Multimodal Sentiment Analysis

  • CAO Hong-fei, CAO Tao
Author information +
文章历史 +

摘要

针对多模态情感分析中模态异质性导致的特征冗余与噪声干扰问题,提出基于多任务联合学习的自适应协同感知模型(AdaSP-MTL)。AdaSP-MTL框架解耦单模态与多模态情感任务,构建异构特征解耦空间以增强模态特异性建模;设计了基于模态间余弦相似度的动态门控融合模块,自适应分配融合权重以抑制跨模态噪声传播;并融合CNN提取的低级局部特征与BiGRU提取的高级时序特征,形成多尺度情感表征。在SIMS、MOSI和MOSEI数据集上的实验评估表明,AdaSP-MTL在情感分类任务上全面优于现有基线模型,验证了其在有效建模模态异质性、抑制噪声干扰以及构建鲁棒情感表征方面的优势。

Abstract

To address the issues of feature redundancy and noise interference caused by modal heterogeneity in multimodal sentiment analysis, this paper proposes an adaptive collaborative perception model based on multi-task joint learning (AdaSP-MTL). The AdaSP-MTL framework decouples unimodal and multimodal sentiment tasks and constructs a heterogeneous feature disentanglement space to enhance modality-specific modeling. A dynamic gated fusion module based on inter-modal cosine similarity is designed to adaptively allocate fusion weights, thereby suppressing cross-modal noise propagation. Furthermore, the model integrates low-level local features extracted by CNNs with high-level sequential features captured by BiGRUs to form multi-scale sentiment representations. Experimental evaluations on the SIMS, MOSI, and MOSEI datasets demonstrate that AdaSP-MTL comprehensively outperforms existing baseline models in sentiment classification tasks, validating its advantages in effectively modeling modal heterogeneity, suppressing noise interference, and constructing robust sentiment representations.

关键词

多模态情感分析 / 多任务联合 / 异质特征 / 自适应融合 / 多尺度特征 / 多模态融合

Key words

multimodal sentiment analysis / multi-task learning / heterogeneous features / adaptive fusion / multi-scale features / multimodal fusion

引用本文

导出引用
曹鸿飞, 曹涛. 协同感知:多任务驱动的自适应多模态情感分析[J]. 电脑与电信. 2026(2): 29-34
CAO Hong-fei, CAO Tao. Collaborative Perception: Multi-Task Driven Adaptive Multimodal Sentiment Analysis[J]. Computer & Telecommunication. 2026(2): 29-34
中图分类号: TP391.1   

参考文献

[1] Holler J,Levinson S C.Multimodal Language Processing in Human Communication[J].Trends in Cognitive Sciences,2019,23(8):639-652.
[2] Miao H,Zhang Y,Wang D,et al.Multi-Output Learning Based on Multimodal GCN and Co-Attention for Image Aesthetics and Emotion Analysis[J].Mathematics,2021,9(12):1437.
[3] Singh P,Srivastava R,Rana K P S,et al.A Multimodal Hierarchical Approach to Speech Emotion Recognition from Audio and Text[J].Knowledge-Based Systems,2021,229:107316.
[4] Zhang Q,Wei Y,Han Z,et al.Multimodal Fusion on Low-Quality Data:A Comprehensive Survey[EB/OL].2024:arXiv:2404.18947.https://arxiv.org/abs/2404.18947
[5] Zhang S,Yang Y,Chen C,et al.Deep Learning-Based Multimodal Emotion Recognition from Audio,Visual,and Text Modalities:A Systematic Review of Recent Advancements and Future Prospects[J].Expert Systems with Applications,2024,237:121692.
[6] Devlin J,Chang M W,Lee K,et al.BERT:Pre-Training of Deep Bidirectional Transformers for Language Understanding[EB/OL].2018:arXiv:1810.04805. https://arxiv.org/abs/1810.04805.
[7] Xu N,Mao W,Chen G.A Co-Memory Network for Multimodal Sentiment Analysis[C]//The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval.ACM,2018:929-932.
[8] Mai S,Hu H,Xing S.Divide,Conquer and Combine:Hierarchical Feature Fusion Network with Local and Global Perspectives for Multimodal Affective Computing[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.Stroudsburg,PA,USA:ACL,2019:481-492.
[9] Zadeh A,Liang P P,Poria S,et al.Multi-Attention Recurrent Network for Human Communication Comprehension[J].Proceedings of the AAAI Conference on Artificial Intelligence,2018,32(1):5642-5649.
[10] Mai S,Hu H,Xu J,et al.Multi-Fusion Residual Memory Network for Multimodal Human Sentiment Comprehension[J].IEEE Transactions on Affective Computing,2022,13(1):320-334.
[11] McFee B,Raffel C,Liang D,et al.Librosa:Audio and Music Signal Analysis in Python[J].Proceedings of the 14th Python in Science Conference,2015:18-24.
[12] Baltrušaitis T,Robinson P,Morency L P.OpenFace:An Open Source Facial Behavior Analysis Toolkit[C]//2016 IEEE Winter Conference on Applications of Computer Vision (WACV).IEEE,2016:1-10.
[13] Zadeh A,Zellers R,Pincus E,et al.MOSI:Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos[EB/OL].2016:arXiv:1606.06259. https://arxiv.org/abs/1606.06259.
[14] Bagher Zadeh A,Liang P P,Poria S,et al.Multimodal Language Analysis in the Wild:CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph[C]//Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1:Long Papers).Association for Computational Linguistics,2018:2236-2246.
[15] Yu W,Xu H,Meng F,et al.CH-SIMS:A Chinese Multimodal Sentiment Analysis Dataset with Fine-Grained Annotation of Modality[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.Association for Computational Linguistics,2020:3718-3727.
[16] Zadeh A,Chen M,Poria S,et al.Tensor Fusion Network for Multimodal Sentiment Analysis[EB/OL].2017:arXiv:1707.07250.https://arxiv.org/abs/1707.07250.
[17] Liu Z,Shen Y,Lakshminarasimhan V B,et al.Efficient Low-Rank Multimodal Fusion with Modality-Specific Factors[EB/OL].2018:arXiv:1806.00064. https://arxiv.org/abs/1806.00064.
[18] A.Zadeh,P.P.Liang,N.Mazumder,S.Poria,E.Cambria,L.-P.Morency,"Memory fusion network for multi-view sequential learning," in Proc.AAAI Conf.Artif.Intell.,vol.32,no.1,Feb.2018.
[19] Koromilas P,Nicolaou M A,Giannakopoulos T,et al.MMATR:A Lightweight Approach for Multimodal Sentiment Analysis Based on Tensor Methods[C]//ICASSP 2023 - 2023 IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP).IEEE,2023:1-5.
[20] Peng J,Wu T,Zhang W,et al.A Fine-Grained Modal Label-Based Multi-Stage Network for Multimodal Sentiment Analysis[J].Expert Systems with Applications,2023,221:119721.
[21] Wang L,Peng J,Zheng C,et al.A Cross Modal Hierarchical Fusion Multimodal Sentiment Analysis Method Based on Multi-Task Learning[J].Information Processing & Management,2024,61(3):103675.
[22] 陈巧红,孙佳锦,漏杨波,等.基于多任务学习与层叠Transformer的多模态情感分析模型[J].浙江大学学报(工学版),2023,57(12):2421-2429.
[23] 王有康,程春玲.基于跨模态单向加权的多模态情感分析模型[J].计算机科学,2025,52(7):226-232.

Accesses

Citation

Detail

段落导航
相关文章

/