为解决多语言文本分类中文本语义表示不全、语义差异大及低资源语言分类准确率低等问题,提升多语言文本分类准确性,提出基于自适应门控特征融合的双通道模型,采用双并行通道处理文本:一个通道通过infoxlm嵌入层进行编码,以保留原始的语言特定特征和语义信息;另一通道则在infoxlm嵌入层编码后集成对抗训练,利用FGM对输入进行扰动,以消除语言特定偏差,增强跨语言语义表示的通用性。两个通道输出的隐藏状态通过门控模块进行动态加权融合,实现两个通道间的信息交互。在多语言毒性检测和情感数据集上的实验表明,所提方法各项评价指标均优于对比模型。该方法能有效提升多语言文本分类准确性及低资源语言的分类效果。
Abstract
To address issues such as incomplete semantic representation, significant semantic variation, and low classification accuracy for low-resource languages in multilingual text classification, thereby enhancing the overall accuracy of multilingual text classification, we propose a dual-channel model based on adaptive gated feature fusion, employing two parallel channels to process text: one channel encodes text via an InfoXLM embedding layer to preserve original language-specific features and semantic information; the other channel integrates adversarial training after InfoXLM embedding, using FGM to perturb inputs to eliminate language-specific biases and enhance the universality of cross-lingual semantic representations. The hidden states from both channels undergo dynamic weighted fusion via a gating module, enabling information exchange between channels. Experiments on multilingual toxicity detection and sentiment datasets demonstrate that the proposed method outperforms comparison models across all evaluation metrics. This approach effectively enhances multilingual text classification accuracy and improves classification performance for low-resource languages.
关键词
多语言文本分类 /
自适应门控融合 /
对抗训练 /
infoxlm /
特征融合
Key words
multilingual text classification /
adaptive gating mechanism /
adversarial training /
InfoXLM /
feature fusion
{{custom_sec.title}}
{{custom_sec.title}}
{{custom_sec.content}}
参考文献
[1] Artetxe M,Schwenk H.Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and beyond[C]//Transactions of the Association for Computational Linguistics.Association for Computational Linguistics,:597-610.
[2] Liu Y,Gu J,Goyal N,et al.Multilingual Denoising Pre-Training for Neural Machine Translation[C]//Transactions of the Association for Computational Linguistics.Association for Computational Linguistics:726-742.
[3] 李文博,高盛祥,张勇丙.基于注意力自适应迁移的零样本跨语言文本分类方法[J].昆明理工大学学报(自然科学版),2025,50(4):95-106.
[4] 于娟,赵慧云,巫邵诚,等.基于句向量加权的跨语言文本分类方法[J].数据分析与知识发现,2025,9(2):39-47.
[5] 崔东虎,崔荣一,赵亚慧.基于无监督对抗训练的跨语言文本分类方法[J].中文信息学报,2023,37(9):55-62.
[6] 李晨,刘纳,郑国风,等.融合双通道特征信息的医疗短文本分类模型[J].现代电子技术,2025,48(13):123-132.
[7] Conneau A,Khandelwal K,Goyal N,et al.Unsupervised Cross-Lingual Representation Learning at Scale[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.Stroudsburg,PA,USA:ACL,2020:8440-8451.
[8] Chi Z,Dong L,Wei F,et al.InfoXLM:An Information-Theoretic Framework for Cross-Lingual Language Model Pre-Training[C]//Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies.Stroudsburg,PA,USA:ACL,2021:3576-3588.
[9] 林楠铠. 基于预训练模型的多语言学习关键技术研究[D].广州:广东工业大学,2024.
[10] Goodfellow I J, Shlens J, Szegedy C.Explaining and Harnessing Adversarial Examples[J].arXiv preprint arXiv:1412.6572, 2014.
[11] Miyato T,Dai A M,Goodfellow I.Adversarial Training Methods for Semi-Supervised Text Classification[EB/OL].2016:arXiv:1605. 07725. https://arxiv.org/abs/1605.07725.
[12] Su D, Zhang H, Chen H, et al.Is robustness the cost of accuracy?-a comprehensive study on the robusnesstof 18 deep image classification models[C]//Proceedings of the European conference on computer vision (ECCV).2018:631-648.
[13] Bahdanau D,Cho K,Bengio Y.Neural Machine Translation by Jointly Learning to Align and Translate[EB/OL].2014:arXiv:1409.0473. https://arxiv.org/abs/1409.0473.
[14] Ma J,Zhao Z,Yi X,et al.Modeling Task Relationships in Multi-Task Learning with Multi-Gate Mixture-of-Experts[C]//Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining.ACM,2018:1930-1939.
[15] FredZhang7,toxi-text-3M.Hugging Face,2023.[EB/OL].Available:https://hf-mirror.com/datasets/FredZhang7/toxi-text-3M.
[16] T.Qiang.A collection of multilingual 3-class sentiments (positive,neutral,negative) dataset.GitHub,2023.[EB/OL].Available:https://github.com/tyqiangz/multilingual-sentiment-datasets.
[17] Wang Z,Lipton Z C,Tsvetkov Y.On Negative Interference in Multilingual Models:Findings and a Meta-Learning Treatment[EB/OL].2020:arXiv:2010.03017.https://arxiv.org/abs/2010.03017 [LinkOut]
[18] Jacob Devlin,Ming-Wei Chang,Kenton Lee,and Kristina Toutanova.2019.BERT:Pre-training of Deep Bidirectional Transformers for Language Understanding.In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies,Volume 1 (Long and Short Papers), pages 4171-4186,Minneapolis,Minnesota.Association for Computational Linguistics.
[19] Zhang Z,Han X,Liu Z,et al.ERNIE:Enhanced Language Representation with Informative Entities[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.Stroudsburg,PA,USA:ACL,2019:1441-1451.
基金
贵州省科技计划项目,项目编号:黔科合基础〔2018〕1082,黔科合基础〔2019〕1159; 贵州省教育厅自然科学研究项目,项目编号:黔教技〔2022〕015,黔教技〔2022〕047,黔教技〔2023〕012,黔教技〔2023〕061,黔教技〔2023〕062; 贵州民族大学基金科研项目,项目编号:GZMUZK〔2023〕YB13